Top Site Reliability Engineer Jobs

Reposted 17 Days AgoSaved
In-Office
2 Locations
Mid level
Mid level
Fintech • Analytics
Maintain SLOs and improve availability, latency, and system health for cloud-hosted services. Design automation and IaC for AWS/Azure, support cloud migration, configure observability (Datadog, CloudWatch, Azure Monitor), participate in on-call rotations and incident response, and partner with development teams to improve reliability, observability, and CI/CD pipelines.
Top Skills: AksAws CloudformationAws CloudwatchAws CodepipelineAws Secrets ManagerAws Well-Architected FrameworkAzure (Services)Azure Architecture CenterAzure Arm TemplatesAzure DevopsAzure MonitorAzure SqlDatadogDockerDynamoDBEc2EksGitIamKmsKubernetesLambdaPythonRdsS3ShellTerraformVpc
Reposted 17 Days AgoSaved
In-Office
Charlotte, NC, USA
Senior level
Senior level
Other
Design, build, and maintain scalable, secure cloud infrastructure and automation. Implement IaC and CI/CD, optimize customer-facing platforms, manage CDNs/DNS/SSL, improve observability and incident response, mentor engineers, and drive infrastructure strategy and reliability improvements.
Top Skills: Akamai CdnAnsibleArgo CdAWSAzureAzure DevopsCdnChefCi/CdCloudFormationComptia Security+DnsGitGCPHTTPHttpsIncident ManagementLinuxMonitoringObservabilityPuppetSsl/TlsTerraformWindows
Reposted 17 Days AgoSaved
Remote or Hybrid
North Carolina, USA
139K-282K Annually
Senior level
139K-282K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Ensure production database services are scalable, resilient, high-performing, and secure. Participate in on-call rotation, monitoring, alerting, incident response, disaster recovery drills, and post-mortems. Automate operational tasks, design reliability/security controls, build observability tooling, capacity forecasts, and collaborate with developers and security to improve hybrid cloud vector and NoSQL database platforms.
Top Skills: AWSAzureCi/CdGCPGitGitJenkinsKubernetesLangchainMilvusMongoDBMySQLNoSQLOpenai ApiOpenshiftPineconePostgresPythonVector DatabasesWeaviate
Reposted 17 Days AgoSaved
In-Office
Chantilly, VA, USA
99K-225K Annually
Senior level
99K-225K Annually
Senior level
Information Technology
Lead SRE responsible for reliability, performance, scalability, observability, automation, and incident response across cloud and air-gapped environments. Drive RCA, capacity planning, reliability standards, and tooling while partnering with DevOps, infrastructure, and security teams to reduce operational risk and support highly available services.
Top Skills: AWSElk StackGrafanaKubernetesLinuxPrometheusPythonTerraformTerragrunt
Reposted 17 Days AgoSaved
In-Office
San Jose, CA, USA
120K-160K Annually
Senior level
120K-160K Annually
Senior level
Cybersecurity
Own reliability, observability, and delivery for a cloud-native, multi-tenant Kubernetes platform. Build CI/CD, GitOps, IaC, and AI-first automation for incident response and LLM infrastructure. Instrument SLOs and tracing, harden secrets and supply-chain security, troubleshoot production issues, and mentor engineers in AI-driven operations.
Top Skills: Agentic OrchestrationCdcCi/CdClaude CodeCloud (Major Providers)CursorData PipelinesDistributed TracingDockerGitopsGoInference GatewaysInfrastructure-As-CodeKubernetesLlmsObservabilityPost-Quantum CryptographyPythonSecrets ManagementService MeshSlosStreamingSupply-Chain SecurityWindsurf
Reposted 23 Days AgoSaved
Easy Apply
Hybrid
Los Angeles, CA, USA
Easy Apply
170K-190K Annually
Senior level
170K-190K Annually
Senior level
AdTech • Big Data • Cloud • Marketing Tech • Software • Analytics
Lead SRE efforts to improve security, reliability, cost efficiency, and observability. Build automation, CI/CD, agentic AI platforms (MCPs), and tooling for capacity planning, incident response, and self-service. Evangelize SecDevOps and zero-trust designs across product and platform teams.
Top Skills: Ai Agentic FrameworksArgocdAWSBashCi/CdDevsecopsDockerEksGoGrafanaKubernetesLinuxLokiMcpNew RelicPrometheusPythonSamTerraformZero Trust
Reposted 23 Days AgoSaved
Hybrid
O'Fallon, MO, USA
96K-163K Annually
Senior level
96K-163K Annually
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead Site Reliability Engineering for Mastercard Business Operations: ensure production readiness, reliability, scalability, and performance of applications. Implement observability, automation, CI/CD, containerization, and cloud infrastructure best practices. Support incident response, capacity planning, troubleshooting, and risk/compliance activities while mentoring developers and promoting developer-run ownership and operational standards.
Top Skills: AWSAzureBashCi/CdContainerizationGCPGoLinuxNetworkingObservabilityOrchestrationPythonUnix
18 Days AgoSaved
In-Office
3 Locations
47K-75K Hourly
Senior level
47K-75K Hourly
Senior level
Financial Services
Deploy, manage, and optimize the Splunk platform and observability tooling. Onboard data sources, build SPL queries and dashboards, monitor platform health, tune performance, respond to incidents/on-call rotations, support Java and .NET applications, perform root cause analysis, and collaborate across engineering, security, and operations to maintain reliable, secure enterprise services.
Top Skills: .NetAnsibleApache Http ServerAppdynamicsApplication Performance ManagementArtifactoryDockerGitGitflowIbm AixIbm Secure Directory ServerIbm Security Access Manager (Isam)JavaJava Enterprise EditionJenkinsJvmKubernetesMicrosoft Windows Server 2008Microsoft Windows Server 2012OpenshiftPythonQuayRed Hat LinuxSearch Processing Language (Spl)SnmpSplunkSplunk Common Information Model (Cim)
Reposted 18 Days AgoSaved
In-Office
Aurora, CO, USA
Expert/Leader
Expert/Leader
Information Technology • Cybersecurity • Defense • Automation
Design, build, and maintain secure, highly available Azure cloud-native platforms and CI/CD pipelines for classified mission systems. Implement Infrastructure as Code, automation, monitoring, and security controls while supporting hybrid Windows environments and platform reliability.
Top Skills: Arm TemplatesAzureAzure ComputeAzure DevopsAzure IdentityAzure Kubernetes ServiceAzure MonitorAzure NetworkingAzure StorageBicepC#Ci/CdDevsecopsDockerGitInfrastructure As CodeKubernetesLog AnalyticsPowershellPythonTerraformWindows Server
Reposted 18 Days AgoSaved
In-Office
Washington, DC, USA
135K-215K Annually
Mid level
135K-215K Annually
Mid level
Aerospace • Defense • Manufacturing
Lead and build the deployment engineering function to operate mission-critical software in accredited, air-gapped, and high-side environments. Own full deployment lifecycle across cloud, on-prem, and disconnected networks; manage Kubernetes/OpenShift and Linux infrastructure; build CI/CD and IaC workflows; integrate security tooling; diagnose and prevent production issues; produce ATO-related artifacts and maintain compliance.
Top Skills: AlertmanagerAWSAzureBashDockerGCPGitlab CiGoGrafanaGroovyHelmJavaJenkinsKubernetesOpenshiftPagerdutyPodmanPrometheusPythonRhelRubyService MeshSplunkTerraform
Reposted 18 Days AgoSaved
Hybrid
Houston, TX, USA
Senior level
Senior level
Hardware • Other • Energy
Maintain and monitor production systems for availability and performance; lead incident response and postmortems; implement observability, alerting, and automated remediation; optimize distributed systems (AKKA.NET) and PostgreSQL; build CI/CD pipelines and infrastructure-as-code.
Top Skills: Akka.NetAWSAzureAzure DevopsAzure PipelinesBashC#DatadogDockerElkGCPGitGithub ActionsGitlabGitlab CiGrafanaKubernetesOpentelemetryPhobosPostgresPowershellPrometheusPythonTerraform
Reposted 18 Days AgoSaved
Remote
USA
180K-210K Annually
Senior level
180K-210K Annually
Senior level
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills: AWSKubernetesPostgresTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 18 Days AgoSaved
In-Office
4 Locations
128K-173K Annually
Senior level
128K-173K Annually
Senior level
Aerospace • Information Technology • Professional Services • Security • Software
Maintain and improve reliability, scalability, and performance of enterprise infrastructure across global sites. Implement automation and infrastructure-as-code, build monitoring and observability, perform RCA and incident response, support patching and RMF changes, integrate new capabilities, and maintain operational documentation and ITIL/ITSM processes to ensure mission-ready, high-availability environments.
Top Skills: AnsibleElkNagiosPowershellPythonScomSolarwindsSplunkTerraform
Reposted 18 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
136K-180K Annually
Senior level
136K-180K Annually
Senior level
Big Data • Energy • Big Data Analytics
The Staff Site Reliability Engineer will lead in designing and maintaining cloud infrastructure on GCP, drive IaC strategy, manage Kubernetes operations, ensure security compliance, and mentor engineers.
Top Skills: BashGoGoogle Cloud PlatformGrafanaKubernetesOpentelemetryPostgresPythonTerraform
Reposted 18 Days AgoSaved
In-Office
Irvine, CA, USA
100K-140K Annually
Junior
100K-140K Annually
Junior
Hardware • Manufacturing
Operate and harden a multi-cloud microservices platform: deploy on Kubernetes, run load/chaos tests, build observability, automate with scripts, define SLO/SLA, ensure security/compliance, participate in incident response, disaster recovery, on-call rotation, and mentor junior team members.
Top Skills: AWSAzureBashGCPGoHpaJavaJvmKubernetesMicroservicesOciPowershellPython
Reposted 18 Days AgoSaved
In-Office
Dallas, TX, USA
Senior level
Senior level
Fintech • Financial Services
Lead SRE technical strategy and architecture for highly available, scalable enterprise platforms. Build automation, observability, and incident response practices; mentor senior engineers; drive capacity planning, production reliability, and adoption of SRE best practices across cloud and on-prem environments.
Top Skills: AnsibleAWSBigQueryChefCloudFormationDatadogDockerElasticsearchElk StackGCPGitlabGoGrafanaJavaJenkinsKafkaKubernetesLinuxMavenPagerdutyPrometheusPrompt EngineeringPuppetPythonRetrieval-Augmented Generation (Rag)Terraform
Reposted 18 Days AgoSaved
Hybrid
2 Locations
130K-160K Annually
Mid level
130K-160K Annually
Mid level
Cloud • Security
Build and operate the production platform (Kubernetes, AWS, IaC, CI/CD, observability), automate self-service deployment, embed security and secrets management, run and modernize on-call, drive cost efficiency, mentor teammates, and maintain runbooks and post-incident reviews.
Top Skills: AWSBashCi/CdClaudeGitGrafanaKubernetesLinuxPrometheusPythonSaltTerraform
Reposted 18 Days AgoSaved
In-Office
2 Locations
157K-239K Annually
Senior level
157K-239K Annually
Senior level
Aerospace • Defense
Own and operate Loft's cloud and hybrid network infrastructure: design architectures, secure connectivity, manage network-as-code (IaC), define SLOs and observability, automate toil, and participate in SRE incident response and platform reliability.
Top Skills: ArgocdCi/CdDnsDockerFirewallingFluxcdGCPGitGitopsGrafanaHybrid ConnectivityK8SKubernetesPeeringSdnSite-To-Site RoutingSoftware-Defined NetworkingTerraformVpcVpn
Reposted 18 Days AgoSaved
In-Office
Aliso Viejo, CA, USA
146K-219K Annually
Senior level
146K-219K Annually
Senior level
Gaming
The role involves ensuring production quality, owning system reliability, and participating in decision-making. Responsibilities include incident response and lifecycle management in cloud gaming technologies.
Top Skills: BashC++ElasticsearchGoIstioJavaKafkaKong Api GatewayKubernetesKumaLinkerdMongoDBMySQLPostgresPythonRedisRust
Reposted 18 Days AgoSaved
Remote
United States
205K-270K Annually
Senior level
205K-270K Annually
Senior level
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills: ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
19 Days AgoSaved
Remote
USA
Senior level
Senior level
Information Technology • Professional Services
Provide senior SRE expertise to improve reliability, scalability, performance, and resilience of a cloud-hosted geospatial platform. Design monitoring/observability, automate deployments, support incident response, optimize capacity and performance, and collaborate across DevSecOps, Kubernetes, database, and support teams in a mission-focused DoD environment.
Top Skills: Alerting ToolsArcgis EnterpriseAw S Cloud OneAWSCi/CdContainerized SystemsEsriInfrastructure-As-CodeKubernetesLinuxLogging ToolsMonitoring ToolsRmfScriptingStig
Reposted 24 Days AgoSaved
Remote or Hybrid
2 Locations
110K-155K Annually
Senior level
110K-155K Annually
Senior level
Information Technology • Insurance • Software
Own and operate production services end-to-end to ensure reliability, scalability, performance, and operational health. Define SLIs/SLOs, perform incident response and root cause analysis, build automation and self-healing, manage production changes, and collaborate with engineering, product, and operations teams to improve system design and observability.
Top Skills: .NetAWSC#Ci/CdInfrastructure As CodeJavaKubernetesLinuxPythonReactRelational DatabasesWindows
Reposted 19 Days AgoSaved
In-Office
4 Locations
208K-269K Annually
Senior level
208K-269K Annually
Senior level
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills: Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Reposted 19 Days AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills: AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
Reposted 19 Days AgoSaved
In-Office
Washington, DC, USA
174K-239K Annually
Senior level
174K-239K Annually
Senior level
Cloud
Design, build, and maintain secure, air-gapped cloud platform services and CI/CD pipelines for Okta Federal. Operate mission-critical infrastructure, monitor SLOs/SLIs, run incident response and POA&M remediation, and support Authority to Operate activities. Advocate SRE/DevOps practices across teams while working autonomously in secure facilities.
Top Skills: Air-Gapped EnvironmentsAmazon CloudwatchAws Transit GatewayAws VpcBgpCi/CdEcs FargateEksGrafanaIpsecPythonSplunkTerraformVpc Endpoints
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account