Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Reposted 8 Days AgoSaved
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
The Staff Engineer will define reliability architecture, automate foundational utilities, develop observability tools, ensure environment integrity, and mentor colleagues.
Top Skills:
AnsibleChefDhcpKubernetesLinuxNtpPxe
Automotive • Hardware • Internet of Things • Mobile • Software • App development • PropTech
Design, implement, and optimize global cloud infrastructure and platforms for an IoT service. Lead platform improvement initiatives, automate infrastructure (IaC/GitOps), ensure observability and security, troubleshoot incidents, mentor SRE team members, and collaborate with executives, architects, and security stakeholders to execute the infrastructure roadmap.
Top Skills:
Active DirectoryArgocdAWSBashDatadogEdge FirewallsGitopsGoGrafanaIacKubernetesLinuxNew RelicPowershellPrometheusPythonSIEMTerraformVpcWindows
Financial Services
Lead and mentor SRE teams to design and operate highly reliable cloud platforms. Drive resiliency reviews, incident leadership, SLO/SLI definition, observability, CI/CD and container practices, and adopt enterprise-authorized AI to accelerate incident triage while ensuring security, auditability, and guardrails.
Top Skills:
.NetAWSCi/CdContainer OrchestrationContainersEnterprise AiJavaMonitoringNetworkingObservabilityPythonSpring BootTelemetry
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and deployment pipelines. Monitor and optimize application availability using observability, SLOs, and telemetry. Automate infrastructure/configuration as code, troubleshoot containers and networking, collaborate across teams, and apply enterprise-authorized AI to accelerate incident triage and reliability improvements while ensuring data sensitivity.
Top Skills:
.NetCi/CdCloudContainersDockerEnterprise AiJavaKubernetesMonitoringNetworkingObservabilityPythonService Level Objectives (Slo)Spring BootTelemetry
Financial Services
Lead design and implementation of observability and reliability for payments services. Define NFRs/SLOs, mentor engineers, drive AI-assisted reliability workflows, reduce toil, ensure auditability/security, and contribute to firm-wide SRE community and tooling.
Top Skills:
AlertingBlack-Box MonitoringEnterprise-Authorized AiObservabilityProduction ReadinessSdlc/ToolchainService Level ObjectivesTelemetry CollectionTesting AutomationWhite-Box Monitoring
Reposted 10 Days AgoSaved
Artificial Intelligence • Software
Operate and maintain on‑prem, air‑gapped Linux infrastructure for US government customers. Design, deploy, and monitor hardware, networking, containers, and orchestration; automate operational tasks; troubleshoot production incidents; collaborate on SLOs and participate in on‑call support. Frequent travel to secure sites and working in classified environments required.
Top Skills:
BashDockerGoJavaJavaScriptKubernetesLinuxOpenshiftPodmanPrometheusPythonRhel
Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
The SRE will ensure the reliability of backend systems, scale Kubernetes-based control planes, and improve automation mechanisms while managing incident processes.
Top Skills:
AWSAzureDockerGCPJavaKubernetesLinuxTerraform
Reposted 10 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Leads reliability standards across Infrastructure Engineering by developing Service Level Objectives, Service Level Indicators, error budgets, observability reporting, and reliability measurement practices. Partners with engineering teams to connect infrastructure performance to application and customer experiences, analyzes distributed-system failure modes, and influences teams through technical guidance, mentoring, and communication with senior leadership.
Top Skills:
Amazon Web ServicesDatadogDistributed SystemsKubernetesObservability PlatformsOn-Premise Infrastructure
Financial Services
Designs, deploys, monitors, and optimizes applications and infrastructure using site reliability engineering practices. Builds infrastructure and configuration as code, develops CI/CD and deployment approaches, improves availability and scalability through SLOs and observability, and resolves complex incidents. The role also applies authorized AI tools for triage and operational analysis, validates recommendations, supports containerized cloud environments, and helps teammates adopt reliability best practices.
Top Skills:
.NetAmazon EcsAnsibleContinuous DeliveryContinuous IntegrationDatadogDockerDynatraceGrafanaJavaKubernetesLinuxPrometheusPythonSplunkSpring BootTerraformWindows
Information Technology • Insurance • Software
Defines enterprise-wide reliability, scalability, performance, and observability standards for critical production services. Leads reliability architecture across AWS, hybrid data centers, and customer-hosted environments; establishes SLIs, SLOs, and error-budget governance; drives fault tolerance and operational sustainability; leads critical incident response and blameless postmortems; and guides teams toward a proactive, engineering-first SRE culture.
Top Skills:
.NetAWSC#Ci/CdError BudgetsInfrastructure-As-CodeJavaKubernetesLinuxObservability FrameworksPythonReactRelational DatabasesSlisSlosWindows
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills:
AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Financial Services
Leads SRE delivery for security-focused applications by enforcing SDLC quality gates, SLOs, operational readiness, resilience, monitoring, and secure CI/CD practices. Oversees incident response, capacity planning, cloud infrastructure, automation, root-cause analysis, and continuous improvement across distributed systems. Partners with Product, Engineering, SRE, and Security leaders while coaching teams on DevOps and reliability principles.
Top Skills:
AppdynamicsAWSCi/CdCloudFormationSplunkTerraform
Financial Services
Maintains and improves application and infrastructure reliability through automation, monitoring, incident response, SLO/SLI management, and infrastructure as code. Designs CI/CD and deployment approaches, troubleshoots cloud, container, networking, and database environments, and proactively reduces operational toil. Collaborates with engineering teams and stakeholders to improve availability, scalability, and operational performance while applying authorized AI tools to accelerate troubleshooting and reliability analysis.
Top Skills:
Amazon EcsAnsibleDatadogDockerDynatraceGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
eCommerce • Legal Tech • Professional Services • Software • Data Privacy
The Site Reliability Engineer will ensure systems run smoothly, work with automation tools, resolve issues, and drive operational improvements.
Top Skills:
AWSAzureCloudFormationDockerGCPGrafanaKubernetesMemcachedNew RelicOpentelemetryPostgresPrometheusPulumiRedisSentryTerraform
Fintech • Software • Financial Services
Lead the SRE function for NinjaTrader’s trading platform, ensuring availability, scalability, performance, and 99.95% uptime. Responsibilities include managing Kubernetes services, resolving production incidents, participating in a 12x7 on-call rotation, automating deployments and operational tasks, designing monitoring and alerting systems, establishing SLIs/SLOs, using Terraform for infrastructure automation, and implementing cloud security and compliance practices. The role also mentors engineers and collaborates cross-functionally on reliable platform delivery.
Top Skills:
AnsibleAWSAzureBashDatadogDockerGCPGithub ActionsGoGrafanaHelmKubernetesPci DssPrometheusPythonSoc 2Terraform
Reposted 13 Days AgoSaved
Financial Services
Lead SRE for AI platforms responsible for driving reliability culture, designing resilient systems, leading incident response, mentoring engineers, establishing SLOs/error budgets, applying AI-assisted operational workflows, and improving observability, CI/CD, container orchestration, and automation across services.
Top Skills:
.NetAgentic AiAiopsBlack-Box MonitoringCi/CdContainer OrchestrationContainersEnterprise Ai ToolingEvent CorrelationJavaKubernetesNetworkingObservabilityPythonRunbook AutomationService Level Objectives (Slos)Spring BootTelemetry CollectionWhite-Box Monitoring
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead and manage an SRE/Platform engineering team to ensure reliability, scalability, and performance of CrowdStrike's cloud-native security platform. Provide technical leadership, incident command, SLO-driven reliability, capacity planning, automation, and mentorship while collaborating with cross-functional teams.
Top Skills:
Apache FlinkApache KafkaAWSAzureElkGCPGoGrafanaIstioJaegerKubernetesLinkerdOpentelemetryPrometheusSplunk
Information Technology • Software • Financial Services • Big Data Analytics
SREs at Citadel focus on optimizing and maintaining system reliability, performance, and automation for investment applications, collaborating closely with teams.
Top Skills:
Ci/CdCSSJavaScriptPythonReactSQL
Artificial Intelligence • Software
The Site Reliability Engineer will maintain high-performance cloud and on-premises services, automate tasks, troubleshoot production issues, and collaborate with product teams.
Top Skills:
AWSAzureBashDockerGCPGoJavaJavaScriptKubernetesLinuxOpenshiftPodmanPrometheusPython
Reposted 14 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
Responsible for managing operations within classified environments, overseeing cloud infrastructure, automating tasks, and ensuring system stability in a high-security setting.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
The Lead Site Reliability Engineer will ensure reliability, scalability, and performance of Mastercard's applications, enhancing operational practices and developer collaboration in a proactive environment.
Top Skills:
Ci/CdDevOpsGoJavaPythonSpring Framework
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Build and maintain automation and reliability for live video distribution across on-prem and cloud. Deploy and manage systems, develop monitoring and automated recovery, troubleshoot complex incidents, coordinate with vendors, document SOPs, support live broadcast components, and participate in L2 on-call rotation.
Top Skills:
AacAc3AnsibleAtscAvcAWSBashChefCloudFormationCmafDockerEksGitHevcHlsJavaScriptJSONKubernetesLinuxMicrosoft Graph ApiMpeg Transport StreamsPythonRistScte104Scte224Scte35SrtSsaiSt2022-7St2110StatmuxTerraformUnixXMLYmlZixi
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills:
Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and SRE practices for critical applications. Build IaC, CI/CD pipelines, observability, and incident response; apply enterprise-authorized AI to accelerate triage while validating outputs and protecting sensitive data.
Top Skills:
.NetAirflowArgo CdAWSAws Step FunctionsCi/CdContainer OrchestrationContainersDatabricksFluxJavaOpentelemetryPythonSpring BootTerraform Enterprise
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results








.png)






.png)












