Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Defense • Manufacturing
Design, deploy, and maintain cloud and on-prem infrastructure with IaC; scale and operate mission-critical services; automate tooling and self-service platforms; collaborate with engineering teams to ensure high availability, reliability, security, and resiliency for vehicle and non-vehicle software systems.
Top Skills:
AWSAzureContainer OrchestrationGCPInfrastructure-As-Code (Iac)KubernetesLinuxNetworking
Aerospace • Other
Build, operate, and scale mission-critical application platforms to accelerate vehicle software delivery. Manage infrastructure as code, improve observability, collaborate with developers, run on-call rotations, conduct blameless postmortems, and reduce performance bottlenecks to support Falcon, Starship, Dragon, and Starlink software lifecycles.
Top Skills:
AnsibleBazelBuckC#C++ClickhouseDockerJavaScriptKubernetesKvmLinuxMakeMySQLPostgresPuppetPythonQemuTerraformVsphere
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Lead reliability and implementation for Kong's Managed Gateways: design and operate multi-cloud, Kubernetes-based systems, own incident response and SLOs, automate CI/CD and IaC, mentor SREs, and drive enterprise customer onboarding and technical implementations.
Top Skills:
AnsibleAWSAzureCassandraCi/CdDatadogElkGCPGoGrafanaIstioKubernetesLinkerdPostgresPrometheusTerraform
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills:
AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Healthtech • Software • Analytics • Business Intelligence
Lead and own reliability for critical backend and distributed systems: design, launch, on-call, incident leadership, SLO/SLI/error budget definition, automation to remove toil, observability improvement, resilience testing, mentoring, and cross-team reliability initiatives for production healthcare workflows.
Top Skills:
AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Hardware • Manufacturing
Lead implementation and operation of microservices on Kubernetes across multi-cloud environments. Build observability, run load/chaos tests, define SLOs/SLA/SLIs, automate with scripts, ensure security/compliance, lead incident response, perform DR planning, mentor teammates, and participate in on-call rotation.
Top Skills:
Application SecurityAWSAzureBashData ProtectionGCPGdprGoHpaIdentity And Access Management (Iam)Iso27001JavaJvmKubernetesMicroservicesNetwork SecurityObservabilityOciPowershellPythonSoc2
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills:
BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Edtech
Maintain and improve site performance, uptime, and scalability. Build monitoring, alerting, runbooks, deployment tooling, and scalable architecture. Troubleshoot across the stack and partner with application teams to deliver reliable production systems.
Top Skills:
AWSBashCC++DockerGCPJavaKubernetesPerlPython
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills:
ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Operate and scale bare-metal Kubernetes clusters at large scale, build control-plane services, operators, and automation for cluster lifecycle. Automate tooling in Go and Python, define SLOs/SLIs, run on-call rotations, troubleshoot incidents, and support customers integrating workloads, storage, and authentication.
Top Skills:
ArgocdCi/CdCluster ApiCniCrdCsiEksFluentbitGitopsGkeGoGrafanaHelmKubeadmKubernetesKubernetes OperatorsPrometheusPython
Software
Operate and scale bare-metal Kubernetes clusters to thousands of nodes, manage incidents and on-call rotation, support customers, collaborate with HPC/datacenter teams, build automation and tooling in Go/Python, create control-plane services/operators, automate cluster lifecycle, and define SLOs/SLIs for platform reliability.
Top Skills:
ArgocdBare-Metal KubernetesCi/CdCluster ApiCniCrdCsiEksFluentbitGkeGoGpuGrafanaHelmKubeadmKubernetesKubernetes OperatorsLinuxPrometheusPython
Other • Energy
Design, build, and operate highly available, scalable systems on Google Cloud. Improve reliability via SLIs/SLOs, capacity planning, DR, and observability. Automate infrastructure with Terraform, improve CI/CD, run on-call rotations, lead incident response and postmortems, and partner with engineering and security teams to reduce toil and optimize cloud cost and performance.
Top Skills:
Azure DevopsBitbucket PipelinesCi/CdDockerGCPGithub ActionsGoGoogle Cloud PlatformIamInfrastructure-As-CodeJavaKubernetesLogsObservability (MetricsPythonService AccountsSlisSlosTerraformTraces)
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for production services in Autodesk GovCloud by building automation, SLO/SLI practices, observability, incident response, resilience testing, and runbooks. Deploy, operate, and improve cloud services while ensuring compliance (FedRAMP) and participating in 24x7 on-call rotations and cross-team collaboration.
Top Skills:
APIsAWSAws GovcloudAzureBashCaching TechnologiesCi/CdCloudwatchContainersDatabasesDatadogDeployment AutomationDistributed SystemsDnsDynatraceGoInfrastructure As CodeJavaKubernetesLoad BalancingMessaging SystemsNetworkingPowershellPythonSplunkStorage Platforms
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for Autodesk GovCloud services by deploying, operating, and automating production systems. Define SLOs/SLIs, build observability and automation, run incident response and on-call rotation, ensure compliance (FedRAMP), perform resilience testing and toil reduction, and collaborate across engineering, security, and platform teams to improve service reliability and operability.
Top Skills:
APIsAWSAws GovcloudAzureBashCaching TechnologiesCi/CdCloudwatchContainersDatabasesDatadogDnsDynatraceFedrampGoIl4Il5Infrastructure As CodeJavaKubernetesLoad BalancingMessaging SystemsNetworkingPowershellPythonSplunkStorage Platforms
Artificial Intelligence • Information Technology • Consulting
Own and operate production infrastructure: manage Kubernetes across regions, maintain IaC and GitOps CI/CD workflows, optimize real-time data pipelines, build observability and alerting, debug incidents, and lead cloud cost and capacity planning for a small engineering team.
Top Skills:
Alerting)Ci/CdGitopsKubernetesMetricsObservability (LoggingTerraform
Artificial Intelligence • Machine Learning • Security • Software
The Senior Staff Site Reliability Engineer will be responsible for ensuring system reliability, debugging issues, mentoring the engineering team, and maintaining infrastructure and CI/CD pipelines.
Top Skills:
AWSDatadogDockerGithub ActionsGrafanaHelmKotlinKubernetesPostgresPrometheusPythonRustTerraformTerragruntTypescript
Healthtech • Professional Services • Pharmaceutical
Lead platform and infrastructure engineering: build full-stack services, design AWS/Terraform infrastructure, manage Kubernetes/Docker, implement CI/CD, improve reliability and security, mentor engineers, and drive compliance for HIPAA and SOC 2.
Top Skills:
ArgocdAWSCloudwatchDatadogDockerGithub ActionsGitopsGrafanaJenkinsKubernetesPrometheusPythonReactTerraformTypescript
Artificial Intelligence • Logistics • Software • Defense
Operate and harden production logistics decision systems for classified and cloud environments. Own availability, monitoring, incident response, CI/CD and IaC automation, and compliance (RMF/STIG/ATO). Partner with engineers and government stakeholders, document runbooks, and support deployments in air-gapped and multi-cloud environments while traveling frequently to customer sites.
Top Skills:
AnsibleAWSAzureAzure Government (Gcc High)Ci/CdDatadogDockerElkGovcloudGrafanaKubernetesLinuxPrometheusTerraform
Software • Financial Services
Own day-to-day AWS and database operations for a serverless production platform, manage backups and disaster recovery, monitor and debug production, lead incident response and on-call, maintain infrastructure-as-code (SST/Pulumi), optimize costs, and mentor the team on operational best practices.
Top Skills:
AWSBashCloudwatchDnsDockerEventbridgeGithub ActionsLambdaLinuxPostgresPulumiPythonS3SqsSstTerraformTlsTypescript
Security • Software • Cybersecurity
Design, build, and operate highly available, fault-tolerant cloud-native systems. Implement observability, automation, CI/CD, and IaC; respond to incidents, run RCAs, and drive reliability improvements.
Top Skills:
AWSAzureBashCi/CdDatadogGCPGoGrafanaInfrastructure As Code (Iac)KubernetesNew RelicOpentelemetryPrometheusPythonSplunkTerraform
Cloud • Software
As a Site Reliability / Gitops Engineer, you will automate operations, develop Infrastructure as Code, maintain core services, and collaborate on service architecture.
Top Skills:
Ci/CdCloud ComputingElasticsearchGrafanaInfrastructure As CodeLinuxPrometheusPython
Fintech
Design, build, and maintain scalable, reliable application infrastructure; automate deployments and recovery; implement observability and monitoring; troubleshoot performance bottlenecks; advise development teams on microservice best practices; produce runbooks; participate in 24x7 on-call rotation and ensure security and compliance.
Top Skills:
AWSCi/CdDockerEncryption ProtocolsGitGoIpJavaJavaScriptKubernetesLinuxObservabilityPythonRubyScripting LanguagesSwarmTcpUdp
Edtech
Maintain and improve site performance, uptime, and scalability. Build monitoring, alerts, runbooks, and deployment tooling. Troubleshoot across the stack, partner with application teams, propose architecture changes, and support on-call operations to ensure reliable product delivery.
Top Skills:
AWSBashCC++DatabasesDockerGCPJavaKubernetesLoad BalancersMessage QueuesMonitoring DashboardsObservability ToolsPerlPythonWeb Servers
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own and improve reliability, performance, and capacity of Lambda's production storage fleet. Build monitoring, alerting, CI/CD, and self-healing automation; investigate and resolve storage incidents; automate operational workflows; collaborate with storage, networking, and release teams; participate in on-call rotation to reduce MTTR and operational toil.
Top Skills:
AlertmanagerAnsibleArgocdBuildkiteCephClushCsi DriversDatadogDell PowerscaleDockerEthtoolGithub ActionsGoGpfsGpudirect StorageGrafanaHelmInfinibandJenkinsKubernetesKustomizeKvm/QemuLinuxLustreMlxlinkNetappNfsNvme-Of/TcpPodmanPrometheusPythonRdmaRoceS3SmbSQLSr-IovSumologicTerraformVastVector DbWeka
Software
Own and improve reliability, performance, and capacity of Lambda's production storage fleet. Build monitoring, alerting, CI/CD, and self-healing automation; troubleshoot storage incidents and low-level I/O/network issues; automate deployments and incident workflows; participate in on-call to reduce MTTR.
Top Skills:
AlertmanagerAnsibleArgocdBuildkiteCephDatadogDockerGithub ActionsGoGpfsGrafanaHelmJenkinsKubernetesKustomizeLinuxLustreNfsNvme-Of/TcpPodmanPrometheusPythonS3SmbSQLSumologicTerraformVector Db
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results



.png)








.png)
.png)






.jpg)




.png)




