Top Site Reliability Engineer Jobs

Reposted One Month AgoSaved
In-Office
Los Angeles, CA, USA
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Defense • Manufacturing
Design, deploy, and maintain cloud and on-prem infrastructure with IaC; scale and operate mission-critical services; automate tooling and self-service platforms; collaborate with engineering teams to ensure high availability, reliability, security, and resiliency for vehicle and non-vehicle software systems.
Top Skills: AWSAzureContainer OrchestrationGCPInfrastructure-As-Code (Iac)KubernetesLinuxNetworking
Reposted One Month AgoSaved
In-Office
Hawthorne, CA, USA
165K-230K Annually
Senior level
165K-230K Annually
Senior level
Aerospace • Other
Build, operate, and scale mission-critical application platforms to accelerate vehicle software delivery. Manage infrastructure as code, improve observability, collaborate with developers, run on-call rotations, conduct blameless postmortems, and reduce performance bottlenecks to support Falcon, Starship, Dragon, and Starlink software lifecycles.
Top Skills: AnsibleBazelBuckC#C++ClickhouseDockerJavaScriptKubernetesKvmLinuxMakeMySQLPostgresPuppetPythonQemuTerraformVsphere
Reposted One Month AgoSaved
In-Office or Remote
Washington, DC, USA
165K-185K Annually
Senior level
165K-185K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Lead reliability and implementation for Kong's Managed Gateways: design and operate multi-cloud, Kubernetes-based systems, own incident response and SLOs, automate CI/CD and IaC, mentor SREs, and drive enterprise customer onboarding and technical implementations.
Top Skills: AnsibleAWSAzureCassandraCi/CdDatadogElkGCPGoGrafanaIstioKubernetesLinkerdPostgresPrometheusTerraform
Reposted One Month AgoSaved
In-Office or Remote
7 Locations
Senior level
Senior level
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills: AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Reposted One Month AgoSaved
Remote or Hybrid
Redmond, WA, USA
120K-150K Annually
Senior level
120K-150K Annually
Senior level
Healthtech • Software • Analytics • Business Intelligence
Lead and own reliability for critical backend and distributed systems: design, launch, on-call, incident leadership, SLO/SLI/error budget definition, automation to remove toil, observability improvement, resilience testing, mentoring, and cross-team reliability initiatives for production healthcare workflows.
Top Skills: AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Reposted One Month AgoSaved
In-Office
Irvine, CA, USA
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Hardware • Manufacturing
Lead implementation and operation of microservices on Kubernetes across multi-cloud environments. Build observability, run load/chaos tests, define SLOs/SLA/SLIs, automate with scripts, ensure security/compliance, lead incident response, perform DR planning, mentor teammates, and participate in on-call rotation.
Top Skills: Application SecurityAWSAzureBashData ProtectionGCPGdprGoHpaIdentity And Access Management (Iam)Iso27001JavaJvmKubernetesMicroservicesNetwork SecurityObservabilityOciPowershellPythonSoc2
Reposted One Month AgoSaved
Remote
US
Senior level
Senior level
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills: BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Reposted One Month AgoSaved
In-Office
Raleigh, NC, USA
Senior level
Senior level
Edtech
Maintain and improve site performance, uptime, and scalability. Build monitoring, alerting, runbooks, deployment tooling, and scalable architecture. Troubleshoot across the stack and partner with application teams to deliver reliable production systems.
Top Skills: AWSBashCC++DockerGCPJavaKubernetesPerlPython
Reposted One Month AgoSaved
Remote
United States
Senior level
Senior level
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills: ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
One Month AgoSaved
Remote or Hybrid
3 Locations
240K-356K Annually
Senior level
240K-356K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Operate and scale bare-metal Kubernetes clusters at large scale, build control-plane services, operators, and automation for cluster lifecycle. Automate tooling in Go and Python, define SLOs/SLIs, run on-call rotations, troubleshoot incidents, and support customers integrating workloads, storage, and authentication.
Top Skills: ArgocdCi/CdCluster ApiCniCrdCsiEksFluentbitGitopsGkeGoGrafanaHelmKubeadmKubernetesKubernetes OperatorsPrometheusPython
One Month AgoSaved
Remote or Hybrid
3 Locations
240K-356K Annually
Senior level
240K-356K Annually
Senior level
Software
Operate and scale bare-metal Kubernetes clusters to thousands of nodes, manage incidents and on-call rotation, support customers, collaborate with HPC/datacenter teams, build automation and tooling in Go/Python, create control-plane services/operators, automate cluster lifecycle, and define SLOs/SLIs for platform reliability.
Top Skills: ArgocdBare-Metal KubernetesCi/CdCluster ApiCniCrdCsiEksFluentbitGkeGoGpuGrafanaHelmKubeadmKubernetesKubernetes OperatorsLinuxPrometheusPython
One Month AgoSaved
In-Office
Houston, TX, USA
Senior level
Senior level
Other • Energy
Design, build, and operate highly available, scalable systems on Google Cloud. Improve reliability via SLIs/SLOs, capacity planning, DR, and observability. Automate infrastructure with Terraform, improve CI/CD, run on-call rotations, lead incident response and postmortems, and partner with engineering and security teams to reduce toil and optimize cloud cost and performance.
Top Skills: Azure DevopsBitbucket PipelinesCi/CdDockerGCPGithub ActionsGoGoogle Cloud PlatformIamInfrastructure-As-CodeJavaKubernetesLogsObservability (MetricsPythonService AccountsSlisSlosTerraformTraces)
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
In-Office or Remote
Plano, TX, USA
117K-209K Annually
Senior level
117K-209K Annually
Senior level
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for production services in Autodesk GovCloud by building automation, SLO/SLI practices, observability, incident response, resilience testing, and runbooks. Deploy, operate, and improve cloud services while ensuring compliance (FedRAMP) and participating in 24x7 on-call rotations and cross-team collaboration.
Top Skills: APIsAWSAws GovcloudAzureBashCaching TechnologiesCi/CdCloudwatchContainersDatabasesDatadogDeployment AutomationDistributed SystemsDnsDynatraceGoInfrastructure As CodeJavaKubernetesLoad BalancingMessaging SystemsNetworkingPowershellPythonSplunkStorage Platforms
Reposted One Month AgoSaved
In-Office
San Francisco, CA, USA
117K-209K Annually
Senior level
117K-209K Annually
Senior level
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for Autodesk GovCloud services by deploying, operating, and automating production systems. Define SLOs/SLIs, build observability and automation, run incident response and on-call rotation, ensure compliance (FedRAMP), perform resilience testing and toil reduction, and collaborate across engineering, security, and platform teams to improve service reliability and operability.
Top Skills: APIsAWSAws GovcloudAzureBashCaching TechnologiesCi/CdCloudwatchContainersDatabasesDatadogDnsDynatraceFedrampGoIl4Il5Infrastructure As CodeJavaKubernetesLoad BalancingMessaging SystemsNetworkingPowershellPythonSplunkStorage Platforms
Reposted One Month AgoSaved
In-Office
New York City, NY, USA
147K-224K Annually
Senior level
147K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Own and operate production infrastructure: manage Kubernetes across regions, maintain IaC and GitOps CI/CD workflows, optimize real-time data pipelines, build observability and alerting, debug incidents, and lead cloud cost and capacity planning for a small engineering team.
Top Skills: Alerting)Ci/CdGitopsKubernetesMetricsObservability (LoggingTerraform
Reposted One Month AgoSaved
In-Office
El Segundo, CA, USA
183K-235K Annually
Senior level
183K-235K Annually
Senior level
Artificial Intelligence • Machine Learning • Security • Software
The Senior Staff Site Reliability Engineer will be responsible for ensuring system reliability, debugging issues, mentoring the engineering team, and maintaining infrastructure and CI/CD pipelines.
Top Skills: AWSDatadogDockerGithub ActionsGrafanaHelmKotlinKubernetesPostgresPrometheusPythonRustTerraformTerragruntTypescript
2 Months AgoSaved
In-Office
Chicago, IL, USA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Healthtech • Professional Services • Pharmaceutical
Lead platform and infrastructure engineering: build full-stack services, design AWS/Terraform infrastructure, manage Kubernetes/Docker, implement CI/CD, improve reliability and security, mentor engineers, and drive compliance for HIPAA and SOC 2.
Top Skills: ArgocdAWSCloudwatchDatadogDockerGithub ActionsGitopsGrafanaJenkinsKubernetesPrometheusPythonReactTerraformTypescript
2 Months AgoSaved
In-Office
Austin, TX, USA
80K-210K Annually
Senior level
80K-210K Annually
Senior level
Artificial Intelligence • Logistics • Software • Defense
Operate and harden production logistics decision systems for classified and cloud environments. Own availability, monitoring, incident response, CI/CD and IaC automation, and compliance (RMF/STIG/ATO). Partner with engineers and government stakeholders, document runbooks, and support deployments in air-gapped and multi-cloud environments while traveling frequently to customer sites.
Top Skills: AnsibleAWSAzureAzure Government (Gcc High)Ci/CdDatadogDockerElkGovcloudGrafanaKubernetesLinuxPrometheusTerraform
2 Months AgoSaved
Remote
US
104K-163K Annually
Senior level
104K-163K Annually
Senior level
Software • Financial Services
Own day-to-day AWS and database operations for a serverless production platform, manage backups and disaster recovery, monitor and debug production, lead incident response and on-call, maintain infrastructure-as-code (SST/Pulumi), optimize costs, and mentor the team on operational best practices.
Top Skills: AWSBashCloudwatchDnsDockerEventbridgeGithub ActionsLambdaLinuxPostgresPulumiPythonS3SqsSstTerraformTlsTypescript
Reposted 2 Months AgoSaved
In-Office
Lehi, UT, USA
125K-145K Annually
Senior level
125K-145K Annually
Senior level
Security • Software • Cybersecurity
Design, build, and operate highly available, fault-tolerant cloud-native systems. Implement observability, automation, CI/CD, and IaC; respond to incidents, run RCAs, and drive reliability improvements.
Top Skills: AWSAzureBashCi/CdDatadogGCPGoGrafanaInfrastructure As Code (Iac)KubernetesNew RelicOpentelemetryPrometheusPythonSplunkTerraform
Reposted One Month AgoSaved
In-Office or Remote
7 Locations
Mid level
Mid level
Cloud • Software
As a Site Reliability / Gitops Engineer, you will automate operations, develop Infrastructure as Code, maintain core services, and collaborate on service architecture.
Top Skills: Ci/CdCloud ComputingElasticsearchGrafanaInfrastructure As CodeLinuxPrometheusPython
2 Months AgoSaved
In-Office
3 Locations
106K-156K Annually
Senior level
106K-156K Annually
Senior level
Fintech
Design, build, and maintain scalable, reliable application infrastructure; automate deployments and recovery; implement observability and monitoring; troubleshoot performance bottlenecks; advise development teams on microservice best practices; produce runbooks; participate in 24x7 on-call rotation and ensure security and compliance.
Top Skills: AWSCi/CdDockerEncryption ProtocolsGitGoIpJavaJavaScriptKubernetesLinuxObservabilityPythonRubyScripting LanguagesSwarmTcpUdp
Reposted 2 Months AgoSaved
In-Office
San Mateo, CA, USA
130K-200K Annually
Senior level
130K-200K Annually
Senior level
Edtech
Maintain and improve site performance, uptime, and scalability. Build monitoring, alerts, runbooks, and deployment tooling. Troubleshoot across the stack, partner with application teams, propose architecture changes, and support on-call operations to ensure reliable product delivery.
Top Skills: AWSBashCC++DatabasesDockerGCPJavaKubernetesLoad BalancersMessage QueuesMonitoring DashboardsObservability ToolsPerlPythonWeb Servers
2 Months AgoSaved
Hybrid
2 Locations
267K-356K Annually
Senior level
267K-356K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own and improve reliability, performance, and capacity of Lambda's production storage fleet. Build monitoring, alerting, CI/CD, and self-healing automation; investigate and resolve storage incidents; automate operational workflows; collaborate with storage, networking, and release teams; participate in on-call rotation to reduce MTTR and operational toil.
Top Skills: AlertmanagerAnsibleArgocdBuildkiteCephClushCsi DriversDatadogDell PowerscaleDockerEthtoolGithub ActionsGoGpfsGpudirect StorageGrafanaHelmInfinibandJenkinsKubernetesKustomizeKvm/QemuLinuxLustreMlxlinkNetappNfsNvme-Of/TcpPodmanPrometheusPythonRdmaRoceS3SmbSQLSr-IovSumologicTerraformVastVector DbWeka
2 Months AgoSaved
Hybrid
2 Locations
267K-356K Annually
Senior level
267K-356K Annually
Senior level
Software
Own and improve reliability, performance, and capacity of Lambda's production storage fleet. Build monitoring, alerting, CI/CD, and self-healing automation; troubleshoot storage incidents and low-level I/O/network issues; automate deployments and incident workflows; participate in on-call to reduce MTTR.
Top Skills: AlertmanagerAnsibleArgocdBuildkiteCephDatadogDockerGithub ActionsGoGpfsGrafanaHelmJenkinsKubernetesKustomizeLinuxLustreNfsNvme-Of/TcpPodmanPrometheusPythonS3SmbSQLSumologicTerraformVector Db
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account