Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Information Technology • Software • Financial Services • Quantitative Trading
The Site Reliability Engineer will provide support and diagnose issues within a real-time, distributed environment, focusing on large-scale application and infrastructure management, with basic required skills in UNIX/Linux, networking, SQL, and scripting languages.
Top Skills:
BashPythonSQLTcp/IpUdpUnix/Linux
Fintech • Information Technology • Financial Services
Own reliability, stability, and performance of low-latency algorithmic and sequencer-based trading platforms. Troubleshoot production issues, manage releases and configurations, improve observability and automation, support client onboarding (FIX/API), conduct performance testing, mentor peers, and drive SRE best practices across teams.
Top Skills:
AWSBashContainerizationFix ProtocolGithub CopilotJavaJenkinsJvmKshLinuxMonitoring/Observability ToolsPythonSQL
Healthtech • Payments • Software
The Senior SRE I will design and maintain automation for infrastructure provisioning, monitor system health, resolve production incidents, and mentor junior SREs, ensuring reliability and operational efficiency across cloud platforms.
Top Skills:
AnsibleAWSAzureBashCloudFormationDatadogDockerGCPGithub ActionsGitlab CiGoGrafanaJavaJenkinsKubernetesPrometheusPythonRubySplunkTerraform
Artificial Intelligence • Marketing Tech • Software
Lead technical reliability initiatives across a multi-cloud, multi-region active-active content platform. Architect and evolve core services, observability and logging, automation and capacity planning. Mentor engineers, drive cross-team reliability projects, define standards (IaC, SLOs, on-call) and proactively improve platform scalability and incident outcomes.
Top Skills:
Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills:
Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Define and scale SRE standards across teams, implement SLOs/SLIs/error budgets, build observability and resiliency patterns, drive automation and AIOps, improve reliability for large-scale Azure cloud systems, and influence engineering and platform teams.
Top Skills:
Ai/MlAiopsAutomationAzureError BudgetsIncident ManagementLogsObservability (MetricsOpentelemetrySlisSlosTracing)
Reposted 24 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills:
AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Financial Services
Designs, implements, monitors, and optimizes cloud-based application infrastructure and reliability. Uses IaC/NaC, observability, SLOs, CI/CD pipelines, and enterprise-authorized AI to prevent and resolve incidents, improve availability and scalability, and mentor peers on SRE best practices.
Top Skills:
Ci/CdCloudContainer OrchestrationContainersContinuous DeliveryContinuous IntegrationEnterprise-Authorized AiInfrastructure As CodeJavaMonitoringNetwork As CodeNetworkingObservabilityPysparkPythonService Level Objectives (Slos)Spring BootTelemetry Collection
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, and operate large-scale, distributed IaaS/PaaS/SaaS systems with automation and monitoring. Provide primary operational support, instrument production KPIs, resolve performance issues, and collaborate across DevOps, Security, and IT teams to scale systems through automation.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Lead and operate reliability for large-scale distributed systems: deploy and monitor IaaS/PaaS/SaaS, automate scaling and deployments, instrument production for KPIs, troubleshoot performance, provide primary operational support, and collaborate with DevOps, Security, and IT Operations to ensure continuous service delivery.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, and operate large-scale distributed IaaS/PaaS/SaaS systems. Implement monitoring, performance tuning, and automation (Ansible/Helm). Provide operational support for networking and storage (Juniper, NFS/Ceph/S3), collaborate with cross-functional teams, and ensure service reliability and continuous improvement.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperNfsOpenstackPaasS3SaaSSecurity+VMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Build, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument performance metrics, scale via automation, troubleshoot complex issues, and collaborate across DevOps, Security, and IT teams.
Top Skills:
AnsibleAzure StackCephHelmIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument production KPIs, resolve complex service issues, scale systems via automation, and collaborate with DevOps, Security, and IT teams to improve reliability and performance.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Lead SRE responsibilities to deploy, automate, monitor, and support large-scale distributed IaaS/PaaS/SaaS systems. Implement automation (Ansible/Helm), instrument production metrics, troubleshoot performance, support storage and virtualization stacks, and collaborate across Core Services, DevOps, Security, and IT Operations to ensure reliable service delivery.
Top Skills:
AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperNfsOpenstackPaasS3SaaSSecurity+VMware
Financial Services
Lead SRE responsible for driving reliability culture, conducting resiliency design reviews, leading incident response, improving service levels with data-driven analytics, adopting AI-assisted SRE workflows, mentoring engineers, and documenting knowledge across the organization.
Top Skills:
.NetAi-Assisted ToolsCi/CdDatadogDockerDynatraceEcsGitlabGrafanaJava Spring BootJenkinsKubernetesObservabilityPrometheusPythonSplunkTelemetryTerraform
Financial Services
Design, implement, and maintain reliable, scalable cloud-native platforms using infrastructure-as-code and CI/CD. Build observability, define SLOs/SLIs, troubleshoot incidents, reduce toil, and collaborate with engineering teams to improve availability and performance.
Top Skills:
AnsibleAWSDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead architecture, design, and modernization of MetLife mainframe platforms. Provide technical leadership for z/OS performance, security (RACF), resilience and BCP automation, partner with stakeholders, tune systems using SMF/RMF telemetry, and drive cross-functional modernization and migration initiatives.
Top Skills:
AnsibleBcpGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRedhat Ansible For Ibm Z CollectionsRmfSafeguarded CopiesSmfWlmZlinux
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills:
Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
Financial Services
Design, implement, monitor, and optimize cloud infrastructure and applications using IaC and CI/CD. Improve availability, reliability, and scalability; use enterprise-authorized AI for incident triage and analysis with strong validation; collaborate across teams; participate in 24x7 on-call incident response.
Top Skills:
.NetAws AlbAws Ec2Aws EksAws NlbAws Route 53Centralized LoggingCi/CdCloudwatchDockerJavaKubernetesKubernetes IngressMetrics/TelemetryMtlsObservabilityPythonSpring BootTerraformTls
Financial Services
Designs, implements, and maintains reliable, scalable infrastructure and deployment pipelines. Automates infrastructure as code, monitors SLOs/SLIs, triages incidents, and improves system reliability collaboratively.
Top Skills:
AnsibleDatadogDockerDynatraceEcsGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Financial Services
Lead architecture and development of production-grade agentic AI and multi-agent systems for reliability engineering. Design autonomous agents, manage AI infrastructure and deployment, ensure observability/security/guardrails, automate remediation, evaluate vendors and models, and produce secure, high-quality production code to improve system reliability and engineering productivity.
Top Skills:
Ai Agent FrameworksAWSCrewaiGenerative AiLanggraphPythonRetrieval-Augmented GenerationTerraformVector Search
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Design, build, and operate reliable, automated CI/CD pipelines and production systems. Monitor availability and performance, conduct incident response and blameless postmortems, drive DevOps automation, provide system design and launch readiness, mentor junior engineers, and collaborate with globally distributed development, operations, and product teams to improve reliability and reduce manual work.
Top Skills:
ArtifactoryBitbucketCC++ChefCi/CdGitGoJavaJenkinsMavenPerlPythonRuby
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Machine Learning • Payments • Security • Software • Financial Services
Senior Site Reliability Engineer responsible for stabilizing and optimizing production environments: platform design, capacity planning, monitoring (SLAs/SLOs), incident response and post-mortems, performance tuning, and mentoring junior engineers to improve reliability and recovery processes.
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead the design, automation, and scaling of global compute infrastructure across data centers, cloud, and on-prem. Operate GitOps with Rancher Fleet/Flux/Helm, build self-healing tooling, own cluster autoscaling and capacity strategy, define SLOs using Datadog, and participate in on-call rotation while mentoring peers.
Top Skills:
AWSContainerdDatadogDockerFluxGCPGitopsGoHelmHpaInfrastructure As Code (Iac)KarpenterKedaKubernetesLinuxNutanixPythonRancher FleetVsphere
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results

























