Top Site Reliability Engineer Jobs

Reposted 8 Days AgoSaved
Hybrid
2 Locations
184K-275K Annually
Senior level
184K-275K Annually
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
The Staff Engineer will define reliability architecture, automate foundational utilities, develop observability tools, ensure environment integrity, and mentor colleagues.
Top Skills: AnsibleChefDhcpKubernetesLinuxNtpPxe
Reposted 8 Days AgoSaved
Hybrid
Oak Brook, IL, USA
103K-193K Annually
Senior level
103K-193K Annually
Senior level
Automotive • Hardware • Internet of Things • Mobile • Software • App development • PropTech
Design, implement, and optimize global cloud infrastructure and platforms for an IoT service. Lead platform improvement initiatives, automate infrastructure (IaC/GitOps), ensure observability and security, troubleshoot incidents, mentor SRE team members, and collaborate with executives, architects, and security stakeholders to execute the infrastructure roadmap.
Top Skills: Active DirectoryArgocdAWSBashDatadogEdge FirewallsGitopsGoGrafanaIacKubernetesLinuxNew RelicPowershellPrometheusPythonSIEMTerraformVpcWindows
Reposted 9 Days AgoSaved
Hybrid
Jersey City, NJ, USA
Senior level
Senior level
Financial Services
Lead and mentor SRE teams to design and operate highly reliable cloud platforms. Drive resiliency reviews, incident leadership, SLO/SLI definition, observability, CI/CD and container practices, and adopt enterprise-authorized AI to accelerate incident triage while ensuring security, auditability, and guardrails.
Top Skills: .NetAWSCi/CdContainer OrchestrationContainersEnterprise AiJavaMonitoringNetworkingObservabilityPythonSpring BootTelemetry
Reposted 9 Days AgoSaved
Hybrid
Chicago, IL, USA
Mid level
Mid level
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and deployment pipelines. Monitor and optimize application availability using observability, SLOs, and telemetry. Automate infrastructure/configuration as code, troubleshoot containers and networking, collaborate across teams, and apply enterprise-authorized AI to accelerate incident triage and reliability improvements while ensuring data sensitivity.
Top Skills: .NetCi/CdCloudContainersDockerEnterprise AiJavaKubernetesMonitoringNetworkingObservabilityPythonService Level Objectives (Slo)Spring BootTelemetry
Reposted 9 Days AgoSaved
Hybrid
Plano, TX, USA
Senior level
Senior level
Financial Services
Lead design and implementation of observability and reliability for payments services. Define NFRs/SLOs, mentor engineers, drive AI-assisted reliability workflows, reduce toil, ensure auditability/security, and contribute to firm-wide SRE community and tooling.
Top Skills: AlertingBlack-Box MonitoringEnterprise-Authorized AiObservabilityProduction ReadinessSdlc/ToolchainService Level ObjectivesTelemetry CollectionTesting AutomationWhite-Box Monitoring
Reposted 10 Days AgoSaved
Hybrid
Washington, DC, USA
125K-185K Annually
Mid level
125K-185K Annually
Mid level
Artificial Intelligence • Software
Operate and maintain on‑prem, air‑gapped Linux infrastructure for US government customers. Design, deploy, and monitor hardware, networking, containers, and orchestration; automate operational tasks; troubleshoot production incidents; collaborate on SLOs and participate in on‑call support. Frequent travel to secure sites and working in classified environments required.
Top Skills: BashDockerGoJavaJavaScriptKubernetesLinuxOpenshiftPodmanPrometheusPythonRhel
Reposted 10 Days AgoSaved
Hybrid
San Francisco, CA, USA
214K-260K Annually
Senior level
214K-260K Annually
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills: AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 10 Days AgoSaved
Easy Apply
Remote or Hybrid
10 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
11 Days AgoSaved
Hybrid
Boston, MA, USA
148K-185K Annually
Senior level
148K-185K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Leads reliability standards across Infrastructure Engineering by developing Service Level Objectives, Service Level Indicators, error budgets, observability reporting, and reliability measurement practices. Partners with engineering teams to connect infrastructure performance to application and customer experiences, analyzes distributed-system failure modes, and influences teams through technical guidance, mentoring, and communication with senior leadership.
Top Skills: Amazon Web ServicesDatadogDistributed SystemsKubernetesObservability PlatformsOn-Premise Infrastructure
11 Days AgoSaved
Hybrid
Jersey City, NJ, USA
Mid level
Mid level
Financial Services
Designs, deploys, monitors, and optimizes applications and infrastructure using site reliability engineering practices. Builds infrastructure and configuration as code, develops CI/CD and deployment approaches, improves availability and scalability through SLOs and observability, and resolves complex incidents. The role also applies authorized AI tools for triage and operational analysis, validates recommendations, supports containerized cloud environments, and helps teammates adopt reliability best practices.
Top Skills: .NetAmazon EcsAnsibleContinuous DeliveryContinuous IntegrationDatadogDockerDynatraceGrafanaJavaKubernetesLinuxPrometheusPythonSplunkSpring BootTerraformWindows
12 Days AgoSaved
Hybrid
Denver, CO, USA
160K-180K Annually
Expert/Leader
160K-180K Annually
Expert/Leader
Information Technology • Insurance • Software
Defines enterprise-wide reliability, scalability, performance, and observability standards for critical production services. Leads reliability architecture across AWS, hybrid data centers, and customer-hosted environments; establishes SLIs, SLOs, and error-budget governance; drives fault tolerance and operational sustainability; leads critical incident response and blameless postmortems; and guides teams toward a proactive, engineering-first SRE culture.
Top Skills: .NetAWSC#Ci/CdError BudgetsInfrastructure-As-CodeJavaKubernetesLinuxObservability FrameworksPythonReactRelational DatabasesSlisSlosWindows
12 Days AgoSaved
Easy Apply
Remote
USA
Easy Apply
241K-270K Annually
Senior level
241K-270K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills: AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
12 Days AgoSaved
Hybrid
Plano, TX, USA
Senior level
Senior level
Financial Services
Leads SRE delivery for security-focused applications by enforcing SDLC quality gates, SLOs, operational readiness, resilience, monitoring, and secure CI/CD practices. Oversees incident response, capacity planning, cloud infrastructure, automation, root-cause analysis, and continuous improvement across distributed systems. Partners with Product, Engineering, SRE, and Security leaders while coaching teams on DevOps and reliability principles.
Top Skills: AppdynamicsAWSCi/CdCloudFormationSplunkTerraform
12 Days AgoSaved
Hybrid
Plano, TX, USA
Mid level
Mid level
Financial Services
Maintains and improves application and infrastructure reliability through automation, monitoring, incident response, SLO/SLI management, and infrastructure as code. Designs CI/CD and deployment approaches, troubleshoots cloud, container, networking, and database environments, and proactively reduces operational toil. Collaborates with engineering teams and stakeholders to improve availability, scalability, and operational performance while applying authorized AI tools to accelerate troubleshooting and reliability analysis.
Top Skills: Amazon EcsAnsibleDatadogDockerDynatraceGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Reposted 12 Days AgoSaved
Remote or Hybrid
4 Locations
175K-175K Annually
Senior level
175K-175K Annually
Senior level
eCommerce • Legal Tech • Professional Services • Software • Data Privacy
The Site Reliability Engineer will ensure systems run smoothly, work with automation tools, resolve issues, and drive operational improvements.
Top Skills: AWSAzureCloudFormationDockerGCPGrafanaKubernetesMemcachedNew RelicOpentelemetryPostgresPrometheusPulumiRedisSentryTerraform
13 Days AgoSaved
Easy Apply
Hybrid
Chicago, IL, USA
Easy Apply
160K-210K Annually
Senior level
160K-210K Annually
Senior level
Fintech • Software • Financial Services
Lead the SRE function for NinjaTrader’s trading platform, ensuring availability, scalability, performance, and 99.95% uptime. Responsibilities include managing Kubernetes services, resolving production incidents, participating in a 12x7 on-call rotation, automating deployments and operational tasks, designing monitoring and alerting systems, establishing SLIs/SLOs, using Terraform for infrastructure automation, and implementing cloud security and compliance practices. The role also mentors engineers and collaborates cross-functionally on reliable platform delivery.
Top Skills: AnsibleAWSAzureBashDatadogDockerGCPGithub ActionsGoGrafanaHelmKubernetesPci DssPrometheusPythonSoc 2Terraform
Reposted 13 Days AgoSaved
Hybrid
Jersey City, NJ, USA
Senior level
Senior level
Financial Services
Lead SRE for AI platforms responsible for driving reliability culture, designing resilient systems, leading incident response, mentoring engineers, establishing SLOs/error budgets, applying AI-assisted operational workflows, and improving observability, CI/CD, container orchestration, and automation across services.
Top Skills: .NetAgentic AiAiopsBlack-Box MonitoringCi/CdContainer OrchestrationContainersEnterprise Ai ToolingEvent CorrelationJavaKubernetesNetworkingObservabilityPythonRunbook AutomationService Level Objectives (Slos)Spring BootTelemetry CollectionWhite-Box Monitoring
Reposted 14 Days AgoSaved
Hybrid
Sunnyvale, CA, USA
140K-215K Annually
Expert/Leader
140K-215K Annually
Expert/Leader
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead and manage an SRE/Platform engineering team to ensure reliability, scalability, and performance of CrowdStrike's cloud-native security platform. Provide technical leadership, incident command, SLO-driven reliability, capacity planning, automation, and mentorship while collaborating with cross-functional teams.
Top Skills: Apache FlinkApache KafkaAWSAzureElkGCPGoGrafanaIstioJaegerKubernetesLinkerdOpentelemetryPrometheusSplunk
Reposted 14 Days AgoSaved
In-Office or Remote
4 Locations
105K-300K Annually
Entry level
105K-300K Annually
Entry level
Information Technology • Software • Financial Services • Big Data Analytics
SREs at Citadel focus on optimizing and maintaining system reliability, performance, and automation for investment applications, collaborating closely with teams.
Top Skills: Ci/CdCSSJavaScriptPythonReactSQL
Reposted 14 Days AgoSaved
Hybrid
Washington, DC, USA
125K-185K Annually
Mid level
125K-185K Annually
Mid level
Artificial Intelligence • Software
The Site Reliability Engineer will maintain high-performance cloud and on-premises services, automate tasks, troubleshoot production issues, and collaborate with product teams.
Top Skills: AWSAzureBashDockerGCPGoJavaJavaScriptKubernetesLinuxOpenshiftPodmanPrometheusPython
Reposted 14 Days AgoSaved
Easy Apply
Remote or Hybrid
Crystal City, VA, USA
Easy Apply
140K-200K Annually
Senior level
140K-200K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Responsible for managing operations within classified environments, overseeing cloud infrastructure, automating tasks, and ensuring system stability in a high-security setting.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 14 Days AgoSaved
Hybrid
O'Fallon, MO, USA
Mid level
Mid level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
The Lead Site Reliability Engineer will ensure reliability, scalability, and performance of Mastercard's applications, enhancing operational practices and developer collaboration in a proactive environment.
Top Skills: Ci/CdDevOpsGoJavaPythonSpring Framework
Reposted 14 Days AgoSaved
Remote or Hybrid
Centennial, CO, USA
110K-145K Annually
Mid level
110K-145K Annually
Mid level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Build and maintain automation and reliability for live video distribution across on-prem and cloud. Deploy and manage systems, develop monitoring and automated recovery, troubleshoot complex incidents, coordinate with vendors, document SOPs, support live broadcast components, and participate in L2 on-call rotation.
Top Skills: AacAc3AnsibleAtscAvcAWSBashChefCloudFormationCmafDockerEksGitHevcHlsJavaScriptJSONKubernetesLinuxMicrosoft Graph ApiMpeg Transport StreamsPythonRistScte104Scte224Scte35SrtSsaiSt2022-7St2110StatmuxTerraformUnixXMLYmlZixi
15 Days AgoSaved
Remote
US
125K-174K Annually
Expert/Leader
125K-174K Annually
Expert/Leader
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills: Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform
15 Days AgoSaved
Hybrid
Wilmington, DE, USA
Mid level
Mid level
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and SRE practices for critical applications. Build IaC, CI/CD pipelines, observability, and incident response; apply enterprise-authorized AI to accelerate triage while validating outputs and protecting sensitive data.
Top Skills: .NetAirflowArgo CdAWSAws Step FunctionsCi/CdContainer OrchestrationContainersDatabricksFluxJavaOpentelemetryPythonSpring BootTerraform Enterprise
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account