Top Site Reliability Engineer Jobs

Reposted 3 Hours AgoSaved
In-Office
6 Locations
125K-350K Annually
Mid level
125K-350K Annually
Mid level
Information Technology • Software • Financial Services • Quantitative Trading
The Site Reliability Engineer will provide support and diagnose issues within a real-time, distributed environment, focusing on large-scale application and infrastructure management, with basic required skills in UNIX/Linux, networking, SQL, and scripting languages.
Top Skills: BashPythonSQLTcp/IpUdpUnix/Linux
Reposted 3 Hours AgoSaved
Hybrid
New York, NY, USA
175K-230K Annually
Senior level
175K-230K Annually
Senior level
Fintech • Information Technology • Financial Services
Own reliability, stability, and performance of low-latency algorithmic and sequencer-based trading platforms. Troubleshoot production issues, manage releases and configurations, improve observability and automation, support client onboarding (FIX/API), conduct performance testing, mentor peers, and drive SRE best practices across teams.
Top Skills: AWSBashContainerizationFix ProtocolGithub CopilotJavaJenkinsJvmKshLinuxMonitoring/Observability ToolsPythonSQL
Reposted YesterdaySaved
In-Office
Louisville, KY, USA
Senior level
Senior level
Healthtech • Payments • Software
The Senior SRE I will design and maintain automation for infrastructure provisioning, monitor system health, resolve production incidents, and mentor junior SREs, ensuring reliability and operational efficiency across cloud platforms.
Top Skills: AnsibleAWSAzureBashCloudFormationDatadogDockerGCPGithub ActionsGitlab CiGoGrafanaJavaJenkinsKubernetesPrometheusPythonRubySplunkTerraform
Reposted YesterdaySaved
Hybrid
Denver, CO, USA
160K-180K Annually
Expert/Leader
160K-180K Annually
Expert/Leader
Information Technology • Insurance • Software
Define and own enterprise reliability, scalability, and performance for production services. Drive architectural standards, observability strategy, SLO/SLI and error-budget governance, lead incident command for high-severity events, and foster a blameless, engineering-first operations culture across cloud, hybrid data centers, and customer-hosted environments.
Top Skills: .NetAWSC#Ci/CdInfrastructure-As-CodeJavaKubernetesLinuxObservabilityPythonReactRelational DatabasesWindows
Reposted YesterdaySaved
Easy Apply
Remote or Hybrid
Ontario, CA, USA
Easy Apply
Senior level
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead technical reliability initiatives across a multi-cloud, multi-region active-active content platform. Architect and evolve core services, observability and logging, automation and capacity planning. Mentor engineers, drive cross-team reliability projects, define standards (IaC, SLOs, on-call) and proactively improve platform scalability and incident outcomes.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
2 Days AgoSaved
Hybrid
Chandler, AZ, USA
119K-224K Annually
Senior level
119K-224K Annually
Senior level
Fintech • Financial Services
Lead and consult on large-scale systems and network infrastructure, resolve complex production issues, drive technical changes, and collaborate with engineering teams to improve reliability and observability across cloud and distributed platforms.
Top Skills: AiopsAirflowArtifactoryAWSAzureBigpandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
2 Days AgoSaved
Hybrid
Columbus, OH, USA
119K-224K Annually
Senior level
119K-224K Annually
Senior level
Fintech • Financial Services
Lead SRE responsible for leading large-scale systems and network infrastructure initiatives, resolving complex production issues, consulting on change/design, and improving observability and hosting platform reliability across cloud and distributed environments while collaborating with technical peers and managers.
Top Skills: AiopsAirflowArtifactoryAWSAzureBig PandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
2 Days AgoSaved
Hybrid
San Francisco, CA, USA
119K-224K Annually
Senior level
119K-224K Annually
Senior level
Fintech • Financial Services
Lead and advise on large-scale systems and network infrastructure initiatives, resolve escalated SRE issues, design technical changes, and collaborate with engineering and management to improve observability, hosting platforms, and production reliability.
Top Skills: AiopsAirflowArtifactoryAWSAzureBigpandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
2 Days AgoSaved
Hybrid
Charlotte, NC, USA
119K-224K Annually
Senior level
119K-224K Annually
Senior level
Fintech • Financial Services
Lead and consult on large-scale systems and network infrastructure planning, drive complex SRE initiatives, analyze escalated technical issues, make technical change decisions, and collaborate with engineering and management to resolve systems support problems and improve reliability.
Top Skills: AiopsAirflowArtifactoryAWSAzureBig PandaElastic ApmElasticsearchGitGradleGrafanaGroovyHarness IoJaegerJenkinsKafkaKibanaKubernetesLinuxLogstashMavenNetcoolOcpPksRemedyServicenowSpinnakerTerraformUdeployUnixVMwareWindowsZipkin
Reposted 2 Days AgoSaved
Easy Apply
Hybrid
New York, NY, USA
Easy Apply
184K-240K Annually
Senior level
184K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead SRE to provide technical leadership for a multi-cloud, multi-region active-active content platform. Define automation and IaC strategy, design core platform and logging/observability systems, establish capacity and performance frameworks, lead cross-team reliability initiatives, improve on-call and incident response, mentor engineers, and proactively address systemic reliability and scaling challenges.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
Reposted 3 Days AgoSaved
In-Office
Berkeley, MO, USA
198K-268K Annually
Expert/Leader
198K-268K Annually
Expert/Leader
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Lead technical strategy and architecture for developer tooling and platforms (GitLab, CI/CD, Jira, Confluence, PostgreSQL, Artifactory, SonarQube). Define SLIs/SLOs, reliability standards, automation, IaC, backups, DR, and security controls. Lead incidents, root-cause analysis, upgrades, migrations, and mentor SREs while partnering with stakeholders to ensure secure, scalable, and supportable platforms. Participate in after-hours escalation as needed.
Top Skills: AnsibleArtifactoryAWSAzure DevopsCi/CdConfluenceContainersDockerGitlabGitlab Ci/CdGCPInfrastructure As CodeJenkinsJIRAKubernetesAzureMonitoring/ObservabilityPostgresRunnersSecrets ManagementSonarqubeVirtualization
Reposted 3 Days AgoSaved
In-Office
Berkeley, MO, USA
99K-171K Annually
Junior
99K-171K Annually
Junior
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Operate and support mission-critical developer platforms (GitLab, CI/CD runners, Jira, Confluence, PostgreSQL, Artifactory, SonarQube). Monitor health, triage incidents, automate operations with IaC/Ansible/scripts, assist cloud/on-prem administration (AWS/Azure/Linux/containers), maintain runbooks, support backups/patching, and collaborate with developers and cybersecurity to apply SRE practices and improve delivery metrics.
Top Skills: AnsibleArtifactoryAWSAzure DevopsBashC#C++ConfluenceContainer PlatformsGitlabGitlab Ci/CdGitlab RunnersInfrastructure As Code (Iac)JavaJenkinsJIRALinuxAzurePostgresPowershellPythonSonarqubeSQLVirtualization
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 3 Days AgoSaved
In-Office or Remote
Minnetonka, MN, USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Define and scale SRE standards across teams, implement SLOs/SLIs/error budgets, build observability and resiliency patterns, drive automation and AIOps, improve reliability for large-scale Azure cloud systems, and influence engineering and platform teams.
Top Skills: Ai/MlAiopsAutomationAzureError BudgetsIncident ManagementLogsObservability (MetricsOpentelemetrySlisSlosTracing)
Reposted 3 Days AgoSaved
Hybrid
San Francisco, CA, USA
214K-260K Annually
Senior level
214K-260K Annually
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills: AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 3 Days AgoSaved
Easy Apply
Remote or Hybrid
7 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills: AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
Reposted 3 Days AgoSaved
Hybrid
O'Fallon, MO, USA
76K-127K Annually
Mid level
76K-127K Annually
Mid level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Ensure reliability, scalability, and performance of Mastercard applications by implementing observability, automation, CI/CD, and cloud infrastructure best practices. Support production readiness, triage incidents, perform root-cause analysis and blameless post-mortems, mentor developers, and drive operational standards, capacity planning, and risk/compliance activities to maximize service availability and customer experience.
Top Skills: AWSAzureBashBitbucketCi/CdContainerizationDynatraceGCPGoJenkinsLinux/UnixOrchestrationPcfPythonSplunkXlr
Reposted 3 Days AgoSaved
Hybrid
2 Locations
197K-246K Annually
Mid level
197K-246K Annually
Mid level
Fintech • Machine Learning • Payments • Software • Financial Services
Lead technical, second-line oversight of SRE and cloud engineering practices. Perform deep-dive risk analyses of cloud architectures, resiliency, CI/CD, observability, and Gen AI integrations. Produce data-driven risk findings, mitigation recommendations, and executive-facing reports while partnering with first-line engineers and leadership to ensure robust controls and operational reliability.
Top Skills: AWSAzureCi/CdCloud-NativeContainerizationDatadogElkGCPGenerative AiKubernetesPagerdutyPrometheusSplunk
4 Days AgoSaved
Hybrid
Houston, TX, USA
Mid level
Mid level
Financial Services
Designs, implements, monitors, and optimizes cloud-based application infrastructure and reliability. Uses IaC/NaC, observability, SLOs, CI/CD pipelines, and enterprise-authorized AI to prevent and resolve incidents, improve availability and scalability, and mentor peers on SRE best practices.
Top Skills: Ci/CdCloudContainer OrchestrationContainersContinuous DeliveryContinuous IntegrationEnterprise-Authorized AiInfrastructure As CodeJavaMonitoringNetwork As CodeNetworkingObservabilityPysparkPythonService Level Objectives (Slos)Spring BootTelemetry Collection
4 Days AgoSaved
In-Office or Remote
Basking Ridge, NJ, USA
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead and scale multidisciplinary engineering teams to deliver cloud-native, AI-enabled enterprise applications and data platforms. Own architecture, modernization, DevOps/MLOps, observability, and AI adoption (RAG, vector DBs, conversational AI). Partner with executives and cross-functional teams to define roadmaps, ensure reliability, and drive engineering excellence and organizational growth.
Top Skills: .NetAWSAzureCi/CdCloud-Native Application ArchitecturesData LakehouseDatabricksDevOpsDistributed Data Processing PlatformsEvent-Driven ArchitecturesFlinkGoogle Cloud PlatformJavaKafkaMicroservicesMlopsOraclePower BIPythonRest ApisSparkSQLSQL Server
Reposted 4 Days AgoSaved
Remote or Hybrid
United States
200K-250K Annually
Senior level
200K-250K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills: Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
5 Days AgoSaved
Hybrid
Merrimack, NH, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, and operate large-scale, distributed IaaS/PaaS/SaaS systems with automation and monitoring. Provide primary operational support, instrument production KPIs, resolve performance issues, and collaborate across DevOps, Security, and IT teams to scale systems through automation.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
5 Days AgoSaved
Hybrid
San Diego, CA, USA
118K-201K Annually
Mid level
118K-201K Annually
Mid level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Lead and operate reliability for large-scale distributed systems: deploy and monitor IaaS/PaaS/SaaS, automate scaling and deployments, instrument production for KPIs, troubleshoot performance, provide primary operational support, and collaborate with DevOps, Security, and IT Operations to ensure continuous service delivery.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
5 Days AgoSaved
Hybrid
Merrimack, NH, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, and operate large-scale distributed IaaS/PaaS/SaaS systems. Implement monitoring, performance tuning, and automation (Ansible/Helm). Provide operational support for networking and storage (Juniper, NFS/Ceph/S3), collaborate with cross-functional teams, and ensure service reliability and continuous improvement.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperNfsOpenstackPaasS3SaaSSecurity+VMware
5 Days AgoSaved
Hybrid
San Diego, CA, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Build, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument performance metrics, scale via automation, troubleshoot complex issues, and collaborate across DevOps, Security, and IT teams.
Top Skills: AnsibleAzure StackCephHelmIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
5 Days AgoSaved
Hybrid
San Diego, CA, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, deploy, automate, monitor, and operate large-scale distributed IaaS/PaaS/SaaS systems. Provide primary operational support, instrument production KPIs, resolve complex service issues, scale systems via automation, and collaborate with DevOps, Security, and IT teams to improve reliability and performance.
Top Skills: AnsibleAzure StackCephHelm ChartsIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSTerraformVMware
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account