Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Artificial Intelligence • Information Technology • Consulting • Cybersecurity
Support and develop data pipelines and telemetry capture for air-gapped cloud environments; collaborate with hardware, firmware, and data science teams; build queries and dashboards; create and standardize deployment solutions; provide operational support and learn new technologies as needed.
Top Skills:
SparkAzureAzure Blob StorageAzure Data ExplorerAzure Data FactoryAzure SynapseC#GitJSONKusto Query Language (Kql)PowershellPublic Key Infrastructure (Pki)PythonSQL
Cloud
Lead design, automation, and operation of large-scale cloud services for federal customers. Improve reliability via SRE practices (SLIs/SLOs, incident response), build automation and self-service platforms using Go/Python, Terraform, and Kubernetes, and ensure compliance with FedRAMP/IL6. Mentor teams, drive reliability initiatives, and modernize deployments with CI/CD and GitOps.
Top Skills:
ArgocdAWSCassandraCi/CdDnsGCPGitopsGoHelmIamKubernetesLoad BalancingMySQLOpensearchPostgresPythonRedisSecrets ManagementTerraformTls
Cloud • Other
Maintain and monitor Rocket.net hosting platform reliability and performance. Provide advanced escalation-level technical support for WordPress VIP customers, troubleshoot Linux-based production environments, web servers, databases, caching, DNS, and networking. Participate in incident response, root cause analysis, automation, documentation, and collaborate with support and engineering teams to improve platform stability.
Top Skills:
ApacheBashCachingCdnCloudflareDatadogDnfDnsHTTPHttpsLinuxMariadbMySQLNedataNginxPhp-FpmRedisSshSsl/TlsWafWordpressYum
3PL: Third Party Logistics
Own and improve uptime for backend services, APIs, workers, and ML pipelines; build monitoring, alerting, auto-remediation, optimize GCP infrastructure and Postgres performance, run on-call rotation and blameless postmortems.
Top Skills:
Apollo ServerBashCi/CdCloud RunCloud SqlDatadogDockerExpressGCPGcsGrafanaNext.JsNode.jsPostgresPrometheusPythonReactTerraformTypescriptZabbix
Financial Services
The Staff Site Reliability Engineer will lead Platform Engineering's SRE efforts by defining technical strategy, overseeing architecture, and enhancing operational excellence through mentorship and governance.
Top Skills:
ArgocdGCPGkeGoKafkaNode.jsPythonTerraform
Artificial Intelligence • Logistics • Robotics • Software
Own reliability across cloud, edge, and on-site deployments. Build observability, monitoring, and alerting. Define incident response and on-call processes, improve deployment workflows, diagnose infra/network/distributed-system issues, and make deployments repeatable and scalable.
Top Skills:
AWSAzureGCPGrafanaKafkaKubernetesLinuxOpentelemetryPrometheusRtspSecure TunnelsVpnWebrtc
Artificial Intelligence • Information Technology • Consulting • Cybersecurity
Design, build, and support telemetry data pipelines and reporting for air-gapped Azure environments. Collaborate with hardware, firmware, and data teams to ingest, transform, query (KQL/SQL), and dashboard telemetry. Create standardized queries and deployment solutions, support air-gapped cloud reporting, and participate in on-call troubleshooting as needed.
Top Skills:
SparkAzureAzure Blob StorageAzure Data ExplorerAzure Data FactoryAzure SynapseC#GitJSONKusto Query Language (Kql)PowershellPublic Key Infrastructure (Pki)PythonSQL
Aerospace • Other
Design, deploy, and operate highly available, sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve performance, and advance deployment, monitoring, and alerting. Collaborate across teams through full software lifecycle to build scalable, operable services for Starlink.
Top Skills:
AlertingApache FlinkApache KafkaSparkBare Metal Compute ClustersC#Continuous IntegrationGoHbaseHdfsIstioJavaKubernetesLinuxMonitoringPythonScalaVersion Control
Aerospace • Other
Design, upgrade, and operate large-scale distributed systems for Starlink. Improve sharding, geo-redundancy, multi-region deployment, monitoring, and performance. Manage petabyte-scale bare-metal clusters and collaborate across teams through the full software lifecycle.
Top Skills:
Apache KafkaSparkC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Aerospace • Other
Design, deploy, and operate sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve deployment/monitoring/alerting, collaborate across teams, and optimize performance throughout the software lifecycle.
Top Skills:
Apache KafkaSparkBare MetalC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Information Technology
Owner of production reliability across cloud and edge: define and drive SLIs/SLOs, build observability (Grafana/Prometheus/Loki/OpenTelemetry), participate in on-call and incident response, encode reliability in infrastructure-as-code (Terraform/OpenTofu), manage Kubernetes clusters, AWS hardening, HA databases, and oversee IoT/edge device fleet operations.
Top Skills:
AWSGitGrafanaIamIotKubernetesLokiNebulaNvidia JetsonOpentelemetryOpentofuPrometheusPyrraSlothStatefulsetsTailscaleTerraformWireguard
Reposted One Month AgoSaved
Fintech • Information Technology • Software • Financial Services
Design, build, and maintain real-time, secure distributed systems and observability UIs/APIs. Implement CI/CD, containerized deployments (Docker/Kubernetes/OpenShift), integrate observability stack (Elasticsearch/Logstash/Grafana), and apply secure coding and API security standards to ensure reliability, performance, and incident automation. Collaborate in Agile teams and explore AI to improve resiliency.
Top Skills:
Agentic AiCi/CdDockerElasticsearchGrafanaJava Spring BootKafkaKubernetesLogstashMariadbNode.jsOauth2OpenshiftReactSecrets Management
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Information Technology • Consulting
Design, build, customize, and support Oracle E-Business Suite and related financial applications. Gather requirements, implement EBS customizations and integrations, optimize performance and SQL, perform testing and upgrades, and provide tier-3 support and issue resolution with Oracle Support.
Top Skills:
Invoice AutomationOracle CloudOracle E-Business SuiteOracle Ebs ApisOracle FinancialsOracle Integration CloudOracle Supplier Portal CloudSQLWebcenter Content Imaging
Financial Services
The Site Reliability Engineer III designs secure, scalable technology solutions, ensures operational resiliency, and collaborates with teams to maintain high availability across environments.
Top Skills:
AutomicAWSAzureBambooBigQueryDockerGitGoogle Cloud PlatformGrafanaJavaJIRAKubernetesLinuxOpentelemetryOraclePostgresPrometheusPythonSplunkUc4Unix
Other
As a Platform Engineer/Dev Ops, you will expand cloud infrastructure, implement monitoring systems, manage databases, and leverage CI/CD tools, working collaboratively with various teams.
Top Skills:
AWSAzureBashDatadogElk StackKubernetesOpentofuPrometheusPythonTerraform
Software
Design, build, and operate multi-account cloud infrastructure using IaC. Automate customer deployments, manage CI/CD, troubleshoot production across infra/data/app layers, and handle networking, security, and compliance for regulated environments while collaborating with platform and professional services teams.
Top Skills:
AirflowAuth0AWSAzureDbtDockerEcsGCPGithub ActionsLlmsOktaPackerPostgresSnowflakeTailscaleTerraformWireguard
Big Data • Cloud • Marketing Tech • Social Impact • Software
The Senior Staff Site Reliability Engineer at LiveRamp will define the SRE strategy, oversee critical automation, and lead operational excellence in a global infrastructure, influencing architectural decisions and mentoring teams.
Top Skills:
Aws)CassandraCircleCICloud Security (GcpDynamoDBGoJenkinsKubernetesPythonScylladbSinglestoreTerraform
Energy
The Senior Site Reliability Engineer improves infrastructure reliability and scalability, partners with various teams, implements IaC and CI/CD, and ensures business continuity through effective BCP/DR planning.
Top Skills:
AWSBashCloudFormationDatadogElkGithub ActionsGitlab CiGoGrafanaJenkinsKubernetesOpensearchPrometheusPythonTerraform
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale distributed systems.
Top Skills:
Cloud ComputingDistributed SystemsHTTPJavaLinux/UnixPerlPythonTcp/IpTls/Ssl
Other
The Sr. Site Reliability Engineer will maintain and administer enterprise systems, troubleshoot operational issues, and develop scripts. This role requires collaboration across teams and participation in project planning and execution.
Top Skills:
AnsibleApacheAzureC#ChefIisJavaJbossPerlPowershellPuppetPythonRubyTomcat
Software
Lead the modernization of AWS cloud infrastructure, implement automation, ensure system reliability, and manage performance with a focus on security and incident response.
Top Skills:
AngularApexAWSC#ElasticacheNew RelicNode.jsNpmPm2PythonRedisShell ScriptingTerraform
Reposted 11 Days AgoSaved
Financial Services
The Senior Site Reliability Engineer will own the operational reliability of developer tooling ecosystems and improve developer productivity through efficient processes and automation.
Top Skills:
.NetBashPowershellPython
Big Data • Analytics • Business Intelligence • Big Data Analytics
Seeking a Site Reliability Engineer to manage AI platform reliability, automate tasks, optimize ML pipelines, and lead incident response in a hybrid engineering role.
Top Skills:
ArgocdBigQueryCloud BuildDockerDvcGithub ActionsGoGrafanaKubeflowKubernetesMlflowPrometheusPub/SubPythonTerraformVertex Ai
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Own reliability, observability, and operational excellence for the Unified Call (911) platform. Design monitoring, alerting, incident response, deployment automation, and dashboards. Improve system resiliency, analyze architecture for operational risks, and build tooling and practices to enable reliable production operations across a Kubernetes-based cloud environment.
Top Skills:
AWSDatadogKafkaKubernetesRabbitMQ
Cloud • Fintech • HR Tech
Responsible for designing, building, automating, and maintaining a secure, highly available Kubernetes-based analytics platform for federal deployments. Improve CI/CD, observability, provisioning (Terraform, Argo CD), troubleshooting, on-call incident response, and collaborate across teams to deliver Workday Prism Analytics in GovCloud.
Top Skills:
SparkArgo CdAWSCi/CdDockerGoGovcloudGrafanaKubernetesObservabilityPrism AnalyticsPrometheusPythonTerraformTracing
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results






























