Top Site Reliability Engineer Jobs

Reposted 21 Days AgoSaved
In-Office
Boston, MA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Lead architecture and build a Kubernetes-based, GitOps-driven platform and self-service datastore offerings. Drive IaC, observability, and platform automation; partner with application teams to diagnose and optimize datastore and messaging performance at scale.
Top Skills: AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesLlm/Agentic Ai ToolingMySQLNew RelicPostgresPulumiTerraform
21 Days AgoSaved
Remote
USA
160K-208K Annually
Senior level
160K-208K Annually
Senior level
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills: AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Reposted 21 Days AgoSaved
Remote or Hybrid
Redmond, WA, USA
120K-150K Annually
Senior level
120K-150K Annually
Senior level
Healthtech • Software • Analytics • Business Intelligence
Lead and own reliability for critical backend and distributed systems: design, launch, on-call, incident leadership, SLO/SLI/error budget definition, automation to remove toil, observability improvement, resilience testing, mentoring, and cross-team reliability initiatives for production healthcare workflows.
Top Skills: AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Reposted 21 Days AgoSaved
In-Office
Raleigh, NC, USA
Senior level
Senior level
Edtech
Maintain and improve site performance, uptime, and scalability. Build monitoring, alerting, runbooks, deployment tooling, and scalable architecture. Troubleshoot across the stack and partner with application teams to deliver reliable production systems.
Top Skills: AWSBashCC++DockerGCPJavaKubernetesPerlPython
22 Days AgoSaved
Remote
United States
Senior level
Senior level
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills: ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
Reposted 22 Days AgoSaved
In-Office or Remote
San Diego, CA, USA
Senior level
Senior level
Information Technology
The Senior Site Reliability Engineer is responsible for architecting reliability strategies, implementing SRE frameworks, mentoring engineers, and ensuring system resilience and performance in government systems.
Top Skills: Cloud ArchitectureDevsecopsGoInfrastructure As Code (Iac)JavaKubernetesLinuxNist 800-53PythonRmf
Reposted 22 Days AgoSaved
In-Office
El Segundo, CA, USA
153K-185K Annually
Senior level
153K-185K Annually
Senior level
Aerospace • Hardware • Software • Biotech • Pharmaceutical • Manufacturing
Lead design, build, and operate mission-critical infrastructure across cloud, on-prem, and spacecraft contexts. Implement IaC, CI/CD, observability, and scalable Kubernetes-based systems; respond to incidents, perform root cause analysis, optimize performance, and collaborate with software and hardware teams. Participate in on-call rotations and occasional travel.
Top Skills: AnsibleArgocdAzureBashCi/CdContainerdDatabasesDockerFirewallsGitopsGpu WorkloadsGrafanaHpcInfluxdbKubernetesLinuxPowershellPrometheusPythonSaltSlurmSubnetsTerraformVpcVpns
Reposted 22 Days AgoSaved
In-Office
Irvine, CA, USA
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Hardware • Manufacturing
Lead implementation and operation of microservices on Kubernetes across multi-cloud environments. Build observability, run load/chaos tests, define SLOs/SLA/SLIs, automate with scripts, ensure security/compliance, lead incident response, perform DR planning, mentor teammates, and participate in on-call rotation.
Top Skills: Application SecurityAWSAzureBashData ProtectionGCPGdprGoHpaIdentity And Access Management (Iam)Iso27001JavaJvmKubernetesMicroservicesNetwork SecurityObservabilityOciPowershellPythonSoc2
Reposted 22 Days AgoSaved
Remote
US
Senior level
Senior level
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills: BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Reposted 22 Days AgoSaved
Remote
United States
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Software
Own and improve platform performance, reliability, and deployment automation. Manage cloud infrastructure, implement IaC, monitor systems with observability tools, provide operational support for distributed applications, and integrate production learnings into development workflows.
Top Skills: Aiops ToolingAws Elastic ContainersAws RdsAws S3Claude CodeClaude CoworkDatadogHarness EngineeringInfrastructure As CodeKubernetesLlmsPrompt EngineeringRigorSplunk
23 Days AgoSaved
In-Office
Atlanta, GA, USA
Senior level
Senior level
Fintech • Information Technology • Payments • Software
Lead design and implementation of a unified enterprise observability platform across Azure, GCP, Kubernetes and hybrid environments. Define observability standards (monitoring, logging, tracing, SLIs/SLOs), build dashboards, integrate with CI/CD and ServiceNow, drive incident prevention/response automation, and promote AI-driven analytics. Influence architecture, mentor engineers, and partner cross-functionally to improve platform reliability and operational readiness.
Top Skills: AksAzureCi/CdDatadogDynatraceGCPGkeGoGoogle Cloud PlatformGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPythonServicenowTerraform
Reposted 24 Days AgoSaved
In-Office
Washington, DC, USA
113K-162K Annually
Senior level
113K-162K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
The role involves operating and scaling Kong's SaaS platform, building automated infrastructure, optimizing multi-region data layers, enhancing observability, and ensuring reliability across services.
Top Skills: ArgocdAWSAzureBashClickhouseDatadogDruidGCPGoGrafanaHelmKubernetesPostgresPrometheusPythonRedisTerraformTerragruntThanos
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 24 Days AgoSaved
In-Office
San Francisco, CA, USA
Senior level
Senior level
Fintech • Payments • Software • Financial Services
Senior SRE responsible for ensuring platform scalability, reliability, and runtime efficiency on AWS. Own CI/CD and GitHub repo workflows, lead incident response and post-mortems, implement observability/monitoring and logging, and collaborate cross-border using bilingual Mandarin and English.
Top Skills: AlertingAWSCi/CdDeployment PipelinesGitGithub ActionsLoggingMonitoringObservabilityScripting
Reposted 24 Days AgoSaved
In-Office
Palo Alto, CA, USA
Senior level
Senior level
Fintech • Payments • Software • Financial Services
Lead Site Reliability Engineer responsible for ensuring platform scalability and uptime on AWS. Own CI/CD and GitHub repository practices, run deployment pipelines, manage incidents and post-mortems, implement observability and logging, and coordinate technical alignment across US and international teams with bilingual communication.
Top Skills: AlertingAWSCi/CdDeployment PipelinesGitGitGithub ActionsLog ManagementMonitoring ToolsObservabilityScripting
Reposted 24 Days AgoSaved
In-Office
San Jose, CA, USA
Senior level
Senior level
Software
Lead architecture, design, and evolution of a global multi-region cloud SRE platform for GPU/AI compute. Author and maintain platform architecture, enforce design invariants, review framework changes, run plugin framework, decide tier placements, coordinate with cloud teams and security, produce pre-flight designs, and shepherd implementations through engineering squads.
Top Skills: BmcDcgmDdnGitopsGpu OperatorInfinibandIpmiKuberayKubernetesKueueLustreMigNcclNetappNvlinkNvme-OfNvswitchPureRayRedfishRoceSlurmSubnet ManagerVastVgpuVolcanoXidZtp
Reposted 24 Days AgoSaved
In-Office
San Jose, CA, USA
Senior level
Senior level
Software
Lead design and implement a global public cloud SRE platform for AI and compute workloads. Own architecture and production engineering for observability, cluster health, remediation, lifecycle, secrets, CI/CD, backup/DR, and automation. Collaborate with cross-functional teams to build scalable, reliable multi-region services and run them in production (on-call).
Top Skills: ArgoAws KmsBmcCosignCrdtDatadogDcgmDdnElasticsearchFluxGcp KmsGoHashicorp VaultHelmInfinibandIpmiJaegerJavaKuberayKubernetesKubernetes Operator (Crd/Controller)KueueKustomizeLokiLustreMimirMtlsNcclNetappNvme-OfOpentelemetryPaxosPrometheusPrometheus QueryPurePythonRaftRayRedfishRoceRustSlurmSQLTempoThanosVastVictoriametricsVolcano
Reposted 24 Days AgoSaved
In-Office
El Segundo, CA, USA
183K-235K Annually
Senior level
183K-235K Annually
Senior level
Artificial Intelligence • Machine Learning • Security • Software
The Senior Staff Site Reliability Engineer will be responsible for ensuring system reliability, debugging issues, mentoring the engineering team, and maintaining infrastructure and CI/CD pipelines.
Top Skills: AWSDatadogDockerGithub ActionsGrafanaHelmKotlinKubernetesPostgresPrometheusPythonRustTerraformTerragruntTypescript
Reposted 24 Days AgoSaved
In-Office
55445, Minneapolis, MN, USA
98K-176K Annually
Senior level
98K-176K Annually
Senior level
eCommerce • Other • Retail
As a Senior Site Reliability Engineer, you will build and support platforms for reliable digital experiences, improve system reliability, and guide technical decisions within the team.
Top Skills: AWSAzureBashDockerFastlyGCPGitGithub ActionsGoKubernetesNext.JsNode.jsReact
Reposted 24 Days AgoSaved
In-Office
New York City, NY, USA
147K-224K Annually
Senior level
147K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Own and operate production infrastructure: manage Kubernetes across regions, maintain IaC and GitOps CI/CD workflows, optimize real-time data pipelines, build observability and alerting, debug incidents, and lead cloud cost and capacity planning for a small engineering team.
Top Skills: Alerting)Ci/CdGitopsKubernetesMetricsObservability (LoggingTerraform
Reposted 24 Days AgoSaved
Remote
United States
141K-208K Annually
Senior level
141K-208K Annually
Senior level
Database • Analytics
This role involves ensuring the reliability and performance of ClickHouse's cloud infrastructure, collaborating with engineering teams, incident management, and driving continuous improvement in service availability.
Top Skills: AnsibleAWSAzureClickhouseDocker SwarmGoGoogle Cloud PlatformKubernetesPuppetPythonTerraform
Reposted 16 Days AgoSaved
In-Office or Remote
7 Locations
Mid level
Mid level
Cloud • Software
As a Site Reliability / Gitops Engineer, you will automate operations, develop Infrastructure as Code, maintain core services, and collaborate on service architecture.
Top Skills: Ci/CdCloud ComputingElasticsearchGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 16 Days AgoSaved
In-Office or Remote
7 Locations
200K-200K Annually
Senior level
200K-200K Annually
Senior level
Cloud • Software
The Senior Site Reliability / Gitops Engineer will drive automation and collaboration within the IS team, enhancing Canonical's IT operations and services while managing infrastructure as code and cloud technologies.
Top Skills: Cloud ComputingDockerElasticsearchGitopsGrafanaIacKubernetesLinuxPrometheusPython
One Month AgoSaved
In-Office
Santa Monica, CA, USA
31-56 Hourly
Junior
31-56 Hourly
Junior
Gaming • Hardware
Entry-level Site Reliability Engineer responsible for monitoring service health, incident response, troubleshooting Kubernetes, networking, DNS, and application issues, building observability (dashboards, alerts, runbooks), automating repetitive tasks, and supporting release reliability and post-incident remediation.
Top Skills: BashCloudContainersDashboardsDnsGitHTTPKubernetesLinuxLoggingMetricsMonitoringPython
9 Months AgoSaved
In-Office
Alexandria, VA, USA
126K-228K Annually
Senior level
126K-228K Annually
Senior level
Information Technology • Software
As a Cloud Site Reliability Engineer, you'll design, implement, and maintain cloud systems across AWS and Azure, ensuring reliability and performance through automation and collaboration with various teams.
Top Skills: Arm TemplatesAWSAws CodepipelineAzureAzure DevopsAzure MonitorBashBicepCloudwatchGithub ActionsGrafanaKubernetesPowershellPrometheusPythonTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account