Top Site Reliability Engineer Jobs

Reposted 21 Days AgoSaved
In-Office
6 Locations
125K-350K Annually
Mid level
125K-350K Annually
Mid level
Information Technology • Software • Financial Services • Quantitative Trading
The Site Reliability Engineer will provide support and diagnose issues within a real-time, distributed environment, focusing on large-scale application and infrastructure management, with basic required skills in UNIX/Linux, networking, SQL, and scripting languages.
Top Skills: BashPythonSQLTcp/IpUdpUnix/Linux
Reposted 21 Days AgoSaved
Hybrid
New York, NY, USA
175K-230K Annually
Senior level
175K-230K Annually
Senior level
Fintech • Information Technology • Financial Services
Own reliability, stability, and performance of low-latency algorithmic and sequencer-based trading platforms. Troubleshoot production issues, manage releases and configurations, improve observability and automation, support client onboarding (FIX/API), conduct performance testing, mentor peers, and drive SRE best practices across teams.
Top Skills: AWSBashContainerizationFix ProtocolGithub CopilotJavaJenkinsJvmKshLinuxMonitoring/Observability ToolsPythonSQL
Reposted 22 Days AgoSaved
In-Office
Louisville, KY, USA
Senior level
Senior level
Healthtech • Payments • Software
The Senior SRE I will design and maintain automation for infrastructure provisioning, monitor system health, resolve production incidents, and mentor junior SREs, ensuring reliability and operational efficiency across cloud platforms.
Top Skills: AnsibleAWSAzureBashCloudFormationDatadogDockerGCPGithub ActionsGitlab CiGoGrafanaJavaJenkinsKubernetesPrometheusPythonRubySplunkTerraform
Reposted 22 Days AgoSaved
Easy Apply
Remote or Hybrid
Ontario, CA, USA
Easy Apply
Senior level
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead technical reliability initiatives across a multi-cloud, multi-region active-active content platform. Architect and evolve core services, observability and logging, automation and capacity planning. Mentor engineers, drive cross-team reliability projects, define standards (IaC, SLOs, on-call) and proactively improve platform scalability and incident outcomes.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Reposted 23 Days AgoSaved
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
184K-240K Annually
Senior level
184K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Reposted 24 Days AgoSaved
In-Office or Remote
Minnetonka, MN, USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Define and scale SRE standards across teams, implement SLOs/SLIs/error budgets, build observability and resiliency patterns, drive automation and AIOps, improve reliability for large-scale Azure cloud systems, and influence engineering and platform teams.
Top Skills: Ai/MlAiopsAutomationAzureError BudgetsIncident ManagementLogsObservability (MetricsOpentelemetrySlisSlosTracing)
Reposted 24 Days AgoSaved
Easy Apply
Remote or Hybrid
7 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills: AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
25 Days AgoSaved
Hybrid
Houston, TX, USA
Mid level
Mid level
Financial Services
Designs, implements, monitors, and optimizes cloud-based application infrastructure and reliability. Uses IaC/NaC, observability, SLOs, CI/CD pipelines, and enterprise-authorized AI to prevent and resolve incidents, improve availability and scalability, and mentor peers on SRE best practices.
Top Skills: Ci/CdCloudContainer OrchestrationContainersContinuous DeliveryContinuous IntegrationEnterprise-Authorized AiInfrastructure As CodeJavaMonitoringNetwork As CodeNetworkingObservabilityPysparkPythonService Level Objectives (Slos)Spring BootTelemetry Collection
26 Days AgoSaved
Hybrid
Merrimack, NH, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, and operate large-scale, distributed IaaS/PaaS/SaaS systems with automation and monitoring. Provide primary operational support, instrument production KPIs, resolve performance issues, and collaborate across DevOps, Security, and IT teams to scale systems through automation.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
26 Days AgoSaved
Hybrid
San Diego, CA, USA
118K-201K Annually
Mid level
118K-201K Annually
Mid level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Lead and operate reliability for large-scale distributed systems: deploy and monitor IaaS/PaaS/SaaS, automate scaling and deployments, instrument production for KPIs, troubleshoot performance, provide primary operational support, and collaborate with DevOps, Security, and IT Operations to ensure continuous service delivery.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
26 Days AgoSaved
Hybrid
Merrimack, NH, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, and operate large-scale distributed IaaS/PaaS/SaaS systems. Implement monitoring, performance tuning, and automation (Ansible/Helm). Provide operational support for networking and storage (Juniper, NFS/Ceph/S3), collaborate with cross-functional teams, and ensure service reliability and continuous improvement.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperNfsOpenstackPaasS3SaaSSecurity+VMware
26 Days AgoSaved
Hybrid
San Diego, CA, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Build, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument performance metrics, scale via automation, troubleshoot complex issues, and collaborate across DevOps, Security, and IT teams.
Top Skills: AnsibleAzure StackCephHelmIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
26 Days AgoSaved
Hybrid
San Diego, CA, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument production KPIs, resolve complex service issues, scale systems via automation, and collaborate with DevOps, Security, and IT teams to improve reliability and performance.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
26 Days AgoSaved
Hybrid
Merrimack, NH, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Lead SRE responsibilities to deploy, automate, monitor, and support large-scale distributed IaaS/PaaS/SaaS systems. Implement automation (Ansible/Helm), instrument production metrics, troubleshoot performance, support storage and virtualization stacks, and collaborate across Core Services, DevOps, Security, and IT Operations to ensure reliable service delivery.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperNfsOpenstackPaasS3SaaSSecurity+VMware
26 Days AgoSaved
Hybrid
Houston, TX, USA
Senior level
Senior level
Financial Services
Lead SRE responsible for driving reliability culture, conducting resiliency design reviews, leading incident response, improving service levels with data-driven analytics, adopting AI-assisted SRE workflows, mentoring engineers, and documenting knowledge across the organization.
Top Skills: .NetAi-Assisted ToolsCi/CdDatadogDockerDynatraceEcsGitlabGrafanaJava Spring BootJenkinsKubernetesObservabilityPrometheusPythonSplunkTelemetryTerraform
26 Days AgoSaved
Hybrid
2 Locations
Mid level
Mid level
Financial Services
Design, implement, and maintain reliable, scalable cloud-native platforms using infrastructure-as-code and CI/CD. Build observability, define SLOs/SLIs, troubleshoot incidents, reduce toil, and collaborate with engineering teams to improve availability and performance.
Top Skills: AnsibleAWSDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Reposted 26 Days AgoSaved
Hybrid
Cary, NC, USA
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead architecture, design, and modernization of MetLife mainframe platforms. Provide technical leadership for z/OS performance, security (RACF), resilience and BCP automation, partner with stakeholders, tune systems using SMF/RMF telemetry, and drive cross-functional modernization and migration initiatives.
Top Skills: AnsibleBcpGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRedhat Ansible For Ibm Z CollectionsRmfSafeguarded CopiesSmfWlmZlinux
27 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
195K-258K Annually
Senior level
195K-258K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills: Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
27 Days AgoSaved
Hybrid
Irvine, CA, USA
Mid level
Mid level
Financial Services
Design, implement, monitor, and optimize cloud infrastructure and applications using IaC and CI/CD. Improve availability, reliability, and scalability; use enterprise-authorized AI for incident triage and analysis with strong validation; collaborate across teams; participate in 24x7 on-call incident response.
Top Skills: .NetAws AlbAws Ec2Aws EksAws NlbAws Route 53Centralized LoggingCi/CdCloudwatchDockerJavaKubernetesKubernetes IngressMetrics/TelemetryMtlsObservabilityPythonSpring BootTerraformTls
27 Days AgoSaved
Hybrid
Plano, TX, USA
Mid level
Mid level
Financial Services
Designs, implements, and maintains reliable, scalable infrastructure and deployment pipelines. Automates infrastructure as code, monitors SLOs/SLIs, triages incidents, and improves system reliability collaboratively.
Top Skills: AnsibleDatadogDockerDynatraceEcsGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
27 Days AgoSaved
Hybrid
Seattle, WA, USA
Senior level
Senior level
Financial Services
Lead architecture and development of production-grade agentic AI and multi-agent systems for reliability engineering. Design autonomous agents, manage AI infrastructure and deployment, ensure observability/security/guardrails, automate remediation, evaluate vendors and models, and produce secure, high-quality production code to improve system reliability and engineering productivity.
Top Skills: Ai Agent FrameworksAWSCrewaiGenerative AiLanggraphPythonRetrieval-Augmented GenerationTerraformVector Search
27 Days AgoSaved
Hybrid
O'Fallon, MO, USA
76K-127K Annually
Mid level
76K-127K Annually
Mid level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Design, build, and operate reliable, automated CI/CD pipelines and production systems. Monitor availability and performance, conduct incident response and blameless postmortems, drive DevOps automation, provide system design and launch readiness, mentor junior engineers, and collaborate with globally distributed development, operations, and product teams to improve reliability and reduce manual work.
Top Skills: ArtifactoryBitbucketCC++ChefCi/CdGitGoJavaJenkinsMavenPerlPythonRuby
Reposted 27 Days AgoSaved
Remote
USA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
28 Days AgoSaved
Hybrid
Cleveland, OH, USA
86K-144K Annually
Senior level
86K-144K Annually
Senior level
Machine Learning • Payments • Security • Software • Financial Services
Senior Site Reliability Engineer responsible for stabilizing and optimizing production environments: platform design, capacity planning, monitoring (SLAs/SLOs), incident response and post-mortems, performance tuning, and mentoring junior engineers to improve reliability and recovery processes.
Reposted 5 Days AgoSaved
Hybrid
Boston, MA, USA
128K-160K Annually
Senior level
128K-160K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills: AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account