Top Site Reliability Engineer Jobs

Reposted 13 Days AgoSaved
In-Office or Remote
3 Locations
100K-125K Annually
Senior level
100K-125K Annually
Senior level
Healthtech • Pet • Biotech
Senior SRE responsible for designing and modernizing CI/CD and deployment systems, automating AWS Serverless infrastructure, improving observability and incident response, enforcing release and security practices, and guiding engineering teams to scale resilient global services.
Top Skills: AuroradbAws CloudformationAws LambdaAzure Entra IdCloudfrontDynamoDBEventbridgeGitGitGithub ActionsMavenOauth2Openid ConnectS3SnsSqsTerraform
Reposted 13 Days AgoSaved
In-Office
Oakland Estates, San Antonio, TX, USA
Senior level
Senior level
Digital Media • Events • Music
Lead and manage a team of SRE/DevOps engineers to ensure reliability, availability, and performance of cloud-based systems. Oversee incident response, operational troubleshooting, process improvements, and cross-team collaboration while mentoring and delegating tasks to meet business objectives.
Top Skills: Cloud Services
Reposted 13 Days AgoSaved
In-Office
Sunnyvale, CA, USA
170K-196K Annually
Senior level
170K-196K Annually
Senior level
Software • Cybersecurity
Senior SRE responsible for ensuring reliability, scalability, and performance of cloud systems (AWS/Azure). Duties include monitoring, on-call support, incident response and RCA, releases and maintenance, security controls, and automation to improve infrastructure efficiency.
Top Skills: AWSAzureAzure DevopsDockerGitlab Ci/CdGoJenkinsKubernetesPowershellPython
Reposted 13 Days AgoSaved
In-Office
Sunnyvale, CA, USA
170K-196K Annually
Senior level
170K-196K Annually
Senior level
Software • Cybersecurity
Drive reliability, scalability, and performance of cloud-based systems on AWS/Azure. Monitor systems, handle on-call production support, lead incident response and root cause analysis, perform releases and hotfixes, implement cloud security controls, and automate infrastructure improvements.
Top Skills: AWSAzureAzure DevopsCloud-NativeDockerGitlab Ci/CdGoJenkinsKubernetesMicroservicesPowershellPython
Reposted 13 Days AgoSaved
In-Office
Reston, VA, USA
133K-238K Annually
Senior level
133K-238K Annually
Senior level
Cloud • Fintech • HR Tech
The Senior Site Reliability Engineer will maintain and optimize the Kubernetes-based platform, ensuring high availability, automating infrastructure, and adhering to security compliance, while troubleshooting and collaborating with development teams.
Top Skills: Argo CdAWSC#Ci/CdGoKubernetesPythonRubyRustTerraform
14 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
140K-288K Annually
Senior level
140K-288K Annually
Senior level
Social Media
Operate, scale, and harden an AWS- and Kubernetes-based platform using GitOps. Build CI/CD and infrastructure-as-code (Terraform/Terragrunt), manage ArgoCD/Helm deployments, improve observability, automate toil reduction, lead incident response and post-incident remediation, and partner with application, security, and platform teams to improve reliability and delivery.
Top Skills: ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonRbacTerraformTerragrunt
Reposted 14 Days AgoSaved
In-Office
Raleigh, NC, USA
119K-196K Annually
Senior level
119K-196K Annually
Senior level
Cloud • Information Technology • Internet of Things • Software • Consulting • Infrastructure as a Service (IaaS) • Automation
Design, build, automate, and operate Red Hat Hybrid OpenShift platforms across cloud and on-prem. Implement GitOps, CI/CD, monitoring, and SRE practices; develop operators and tooling; lead incident response, on-call duties, and postmortems; mentor peers and improve platform reliability and self-service.
Top Skills: ArgocdAWSAzureCatchpointCi/CdDatadogDnsFedoraGitlabGitopsGoGoGCPGrafanaHttp/TlsKubernetesKubernetes OperatorsLdapLinuxOpenshiftOpenshift PipelinesOperator SdkPrometheusPythonRhelSplunkSplunk ImTcp/IpTekton
15 Days AgoSaved
In-Office
City Point, City of Boston, MA, USA
130K-140K Annually
Senior level
130K-140K Annually
Senior level
Fintech • Software
Senior SRE responsible for ensuring reliability, scalability, and performance of production systems. Investigates and resolves incidents, works with R&D on defects, manages deployments and change validation, builds monitoring and diagnostic tooling to improve MTTA/MTTD/MTTR, configures observability platforms, participates in on-call rotation, and performs risk assessments and production readiness validation.
Top Skills: Ai ToolsAkamaiAmqAnsibleAWSAzureBashCdnCloudflareCloudwatchDatadogDatastreamDnsDockerDynatraceEfsEksHTTPHttpsInterconnectJavaJmsKubernetesMongoDBOraclePostgresPrometheusPythonRabbitMQRedshiftS3SplunkTcpTerraformUdpUnix/LinuxWafZabbix
15 Days AgoSaved
In-Office
Tyson's Corner, VA, USA
165K-200K Annually
Senior level
165K-200K Annually
Senior level
Fintech
Lead the reliability, monitoring, and scaling of Kubernetes clusters and cloud infrastructure. Deploy and maintain containerized microservices, implement IaC (Terraform/CloudFormation), automate operational tasks, manage incidents, and ensure security and compliance across environments. Collaborate with development and operations teams to optimize performance and availability.
Top Skills: Amazon S3AnsibleApache MesosAWSAzureC/C++CephCloudFormationDockerGCPHdfsHelmJavaJavaScriptJenkinsKubernetesLinuxNfsPostgresPythonRubyTerraformYarn
Reposted YesterdaySaved
Remote or Hybrid
4 Locations
165K-330K Annually
Mid level
165K-330K Annually
Mid level
Software
As a Site Reliability Engineer, you'll build and maintain infrastructure for ML models, automate processes, and collaborate cross-functionally.
Top Skills: Circle CiCloudFormationElk StackGithub ActionsGitlab CiGrafanaJenkinsKubernetesOpentelemetryPrometheusPulumiTerraform
Reposted 15 Days AgoSaved
Hybrid
Bellevue, WA, USA
120K-150K Annually
Senior level
120K-150K Annually
Senior level
Healthtech • Software • Analytics • Business Intelligence
Senior SRE responsible for designing, building, and operating reliable, scalable distributed systems; owning production reliability (SLOs/SLIs, incident response, MTTR reduction); automating toil with software and platform tooling; driving observability, capacity planning, and cross-team reliability improvements; mentoring engineers and running blameless postmortems.
Top Skills: AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Reposted 15 Days AgoSaved
Remote
United States
82K-229K Annually
Senior level
82K-229K Annually
Senior level
Cloud • Software
Design, implement, and support Kubernetes and compute platforms in a private cloud. Oversee architecture and standardization across hardware, OS, and cloud orchestration.
Top Skills: AnsibleBashCi/CdHelmKubernetesLinuxOpenstackPythonTerraformUbuntu
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 15 Days AgoSaved
In-Office
San Francisco, CA, USA
210K-240K Annually
Senior level
210K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
The Senior Site Reliability Engineer will design and maintain scalable infrastructure, improve system reliability, manage CI/CD pipelines, and collaborate across teams for operational excellence.
Top Skills: AnsibleArgocdAWSBashDatadogDockerElkGithub ActionsGrafanaKubernetesLinuxOpentelemetryPrometheusPythonTerraform
Reposted 15 Days AgoSaved
Hybrid
2 Locations
150K-175K Annually
Expert/Leader
150K-175K Annually
Expert/Leader
Artificial Intelligence • Machine Learning • Natural Language Processing • Software
As a Site Reliability Engineer, you'll enhance the performance and reliability of infrastructure and products by collaborating with engineering teams, automating configurations, and implementing monitoring systems.
Top Skills: AWSGoKubernetesPythonTerraform
Reposted 15 Days AgoSaved
In-Office
Sunnyvale, CA, USA
170K-196K Annually
Senior level
170K-196K Annually
Senior level
Software • Cybersecurity
The Senior Site Reliability Engineer will ensure the reliability and performance of cloud systems, manage AWS/Azure infrastructure, optimize performance, lead incident responses, and implement security best practices.
Top Skills: AWSAzureAzure DevopsDockerGitlab Ci/CdGoJenkinsKubernetesPowershellPython
Reposted 15 Days AgoSaved
In-Office
2 Locations
Senior level
Senior level
Financial Services
Own reliability and scalability of on-prem observability platforms (ELK, Grafana); handle production escalations, capacity planning, SLOs, onboarding, automation, IaC (Terraform/Helm/Ansible), upgrades, security hardening, and platform modernization.
Top Skills: AnsibleApm InstrumentationBashBeatsChefElasticsearchElk StackFluent BitFluentdGrafanaHelmKibanaLinuxLogstashNew RelicOpentelemetryPrometheusPuppetPythonShell Scripting/Linux ShellSolarwindsTerraform
Reposted 15 Days AgoSaved
Remote
USA
113K-176K Annually
Senior level
113K-176K Annually
Senior level
Other • Social Impact
As a Senior Site Reliability Engineer, you will manage and improve Wikimedia's infrastructure, handle operational tasks, automate processes, and provide mentorship while participating in a 24/7 on-call rotation.
Top Skills: AnsibleBashDebianGoGrafanaHhvmKubernetesMemcachedPHPPrometheusPuppetPythonRedisRuby
Reposted 15 Days AgoSaved
In-Office
San Francisco, CA, USA
135K-159K Annually
Senior level
135K-159K Annually
Senior level
Big Data • Cloud • Marketing Tech • Social Impact • Software
The Senior Site Reliability Engineer will support global product deployments, provide 24/7 operational support, maintain CI/CD tooling, and optimize system performance. They will utilize their SRE practices and leadership abilities to improve product reliability and guide other engineers.
Top Skills: AWSCassandraCircleCIDynamoDBGCPGoJenkinsKubernetesNosql DatabasesPythonScylladbSinglestore DbTerraform
Reposted 15 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
167K-197K Annually
Senior level
167K-197K Annually
Senior level
Big Data • Cloud • Marketing Tech • Social Impact • Software
As a Senior Site Reliability Engineer, you will manage global product deployments, provide operational support, enhance CI/CD tooling, and optimize system performance, collaborating closely with distributed engineering teams.
Top Skills: AWSCassandraCircleCIDynamoDBGCPGoJenkinsKubernetesNosql DatabasesPythonScylladbSinglestore DbTerraform
16 Days AgoSaved
Remote
United States
108K-136K Annually
Senior level
108K-136K Annually
Senior level
Big Data • Cloud • Hardware • Software • App development
Senior Site Reliability Engineer responsible for improving reliability and performance of wwt.com. Build monitoring/alerting, lead incident response and root cause analysis, automate operational tasks, run load testing, review penetration test results, mentor engineers, and collaborate with developers to ensure scalable, resilient systems.
Top Skills: AngularGoJavaNode.jsPythonReactVue
16 Days AgoSaved
Remote
United States
108K-136K Annually
Senior level
108K-136K Annually
Senior level
Cloud • Information Technology • Productivity • Security • Software
Senior Site Reliability Engineer responsible for improving reliability, performance, and scalability of wwt.com. Develop monitoring, alerting, and automation; lead incident response and root cause analysis; perform load testing and security vulnerability remediation; collaborate with developers and QA; write code to automate operations; mentor engineers and influence technical direction.
Top Skills: AngularGoJavaNode.jsPythonReactVue
16 Days AgoSaved
In-Office
Palo Alto, CA, USA
165K-280K Annually
Senior level
165K-280K Annually
Senior level
Aerospace • Other
Design, deploy, and operate highly available, sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve performance, and advance deployment, monitoring, and alerting. Collaborate across teams through full software lifecycle to build scalable, operable services for Starlink.
Top Skills: AlertingApache FlinkApache KafkaSparkBare Metal Compute ClustersC#Continuous IntegrationGoHbaseHdfsIstioJavaKubernetesLinuxMonitoringPythonScalaVersion Control
16 Days AgoSaved
In-Office
Redmond, WA, USA
165K-270K Annually
Senior level
165K-270K Annually
Senior level
Aerospace • Other
Design, deploy, and operate sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve deployment/monitoring/alerting, collaborate across teams, and optimize performance throughout the software lifecycle.
Top Skills: Apache KafkaSparkBare MetalC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
16 Days AgoSaved
In-Office
Hawthorne, CA, USA
165K-265K Annually
Senior level
165K-265K Annually
Senior level
Aerospace • Other
Design, upgrade, and operate large-scale distributed systems for Starlink. Improve sharding, geo-redundancy, multi-region deployment, monitoring, and performance. Manage petabyte-scale bare-metal clusters and collaborate across teams through the full software lifecycle.
Top Skills: Apache KafkaSparkC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Reposted 17 Days AgoSaved
In-Office
Santa Clara, CA, USA
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Software
Lead the modernization of AWS cloud infrastructure, implement automation, ensure system reliability, and manage performance with a focus on security and incident response.
Top Skills: AngularApexAWSC#ElasticacheNew RelicNode.jsNpmPm2PythonRedisShell ScriptingTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account