Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Fintech • Financial Services
Design, build, and maintain reliable, scalable virtual desktop infrastructure (VDI) and supporting platforms. Lead incident response, automate deployments and operations with IaC and CI/CD, implement secure configurations, monitor system health, collaborate cross-functionally, and drive continuous improvement and operational excellence.
Top Skills:
Active DirectoryAnsibleArm/BicepAzure DevopsCitrix CloudCitrix GatewayCvadDnsDscGithub ActionsGitlab CiGposJenkinsPowershellSsl/Tls CertificatesTerraformVdi Profile ManagementWindows 11 Multi-SessionWindows Server
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills:
ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Payments
Own platform reliability, security, automation, uptime, and incident response for payment-processing systems. Build and evolve AWS infrastructure with Terraform, manage infrastructure as code across cloud and physical environments, containerize legacy services, administer Linux servers, and embed PCI DSS compliance through patching, vulnerability management, logging, and monitoring. Provide technical leadership and support secure infrastructure for analytics and BI systems.
Top Skills:
AnsibleAWSContainersLinuxOpenvoxPci DssPuppetTerraform
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills:
AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
Hardware • Information Technology • Other • Software • Analytics
Configure, optimize, and support AWS and Azure cloud infrastructure, especially AWS GovCloud environments supporting FedRAMP-compliant services. Improve availability, performance, security, scalability, and capacity; manage monitoring, application operations, incident response, runbooks, architecture documentation, and compliance. Collaborate with development teams and international stakeholders while using scripting, infrastructure-as-code, container, CI/CD, and observability tools.
Top Skills:
AnsibleAWSAws GovcloudAzureBashEcsEksFedrampGitGrafanaJenkinsJIRAKubernetesPerlPowershellPrtgPythonSumo LogicTerraform
Hardware • Information Technology • Other • Software • Analytics
Design, build, and maintain scalable cloud infrastructure across AWS GovCloud and Azure. Develop automated deployment, monitoring, observability, incident response, and resilience testing systems. Establish SLOs, manage error budgets, troubleshoot application operations, and enforce FedRAMP security practices. Collaborate with architecture and software engineering teams to improve availability, performance, and automation across mission-critical SaaS and PaaS platforms.
Top Skills:
Active DirectoryAnsibleAWSAws GovcloudAws SdkAzureBashDockerEcsEksFedrampGrafanaJira Service ManagementKubernetesNoSQLPerlPowershellPrtgPythonRdbmsSAMLSumoTerraform
5 Days AgoSaved
Financial Services
Lead application support and SRE activities for mission-critical systems, improving reliability, observability, resilience, and operational efficiency. Responsibilities include resolving production incidents, managing incident and problem processes, conducting root-cause reviews, supporting releases and disaster recovery testing, maintaining operational documentation, optimizing alerts, automating workflows, and strengthening risk and control practices. The role partners with engineering, infrastructure, operations, and global teams in a highly regulated financial services environment.
Top Skills:
ActivemqAmazon Rds AuroraAutosysAws Ec2Aws IamAws LambdaAws S3Aws SqsBashCicsCobolDb2Db2 Stored ProceduresDynatraceFile-AidGrafanaIbm MqJavaScriptJclJIRAKafkaLinuxOpenshiftOracleOracle AqPerlPostgresPythonRabbitMQRubySeleniumServicenowShell ScriptingSnowflakeSplunkSpufiSQLWindows
Mining Operations
Leads reliability improvements for steel-producing equipment and facilities, reducing downtime through predictive and preventive maintenance programs. Uses Maximo and equipment databases, analyzes failure data, and supports criticality assessments, spare-parts analysis, root cause analysis, and FMEA. Provides troubleshooting and reliability expertise, collaborates on new and modified installations, updates engineering standards, and develops best practices through periodic travel to other company sites.
Top Skills:
CmmsElectrical SystemsFmeaHydraulicsIbm Maximo EamMechanical SystemsPneumaticsReliability-Centered Maintenance
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills:
Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
Gaming • Mobile
Enhance production-system stability, performance, scalability, and reliability through observability, automation, Kubernetes operations, and AWS infrastructure management. Define and monitor SLIs, SLOs, and SLAs; manage incidents, on-call response, troubleshooting, root-cause analysis, and post-incident remediation. Improve CI/CD production readiness, secrets management, and application resilience while documenting runbooks and operational procedures.
Top Skills:
.NetAmazon Ec2Amazon EksAmazon Route 53Amazon S3Argo CdAWSAws IamBashGithub ActionsGitlab Ci/CdGraylogHashicorp VaultHelmKubernetesNew RelicPackerPythonRancherTerraform
Energy
Build and support scalable, resilient cloud-native platforms and applications. Responsibilities include reliability engineering, observability, incident response, platform engineering, cloud infrastructure automation, Infrastructure as Code, CI/CD enablement, disaster recovery, chaos engineering, and AI-driven operations. The role leads reliability improvements, develops self-service tooling, partners cross-functionally on platform modernization, and mentors junior engineers.
Top Skills:
APIsAWSAzureBashChaos EngineeringCi/CdCloudFormationDatadogDistributed SystemsDockerGitGitopsGrafanaHelmInfrastructure As CodeKubernetesLinuxMicroservicesNew RelicOpentelemetryPrometheusPythonSplunkTerraform
Digital Media • Gaming • News + Entertainment • Sports
Lead multiple SRE teams to ensure reliability, scalability, and security across cloud and on-prem systems. Drive observability, automation, CI/CD, infrastructure-as-code, and operational excellence; set strategy, manage resources, mentor leaders, and influence stakeholders for commerce platforms.
Top Skills:
Ai/MlAnsibleAWSAzureCi/CdCloudFormationGCPGitlabHarnessInfrastructure-As-CodeKubernetesServerlessTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Information Technology • Professional Services • Software • Consulting
Design, deploy, and maintain highly available production systems across cloud environments. Build automation and reliability tooling, manage Kubernetes and Docker workloads, implement Infrastructure as Code and CI/CD pipelines, and develop monitoring and observability using OpenTelemetry, Prometheus, and Grafana. Troubleshoot complex infrastructure and application issues, participate in incident response and root-cause analysis, improve system performance and capacity, and establish SRE reliability practices.
Top Skills:
AnsibleAWSAzureCi/CdDockerGCPGoGrafanaJavaKubernetesLinuxOpentelemetryPrometheusPythonRustTerraformUnix
Reposted 10 Days AgoSaved
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills:
AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Artificial Intelligence • Software
Design, build, and scale control- and data-plane infrastructure for distributed AI workloads. Improve reliability, performance, scheduling, and observability for Ray clusters across cloud and on-prem environments. Support accelerator integration, container image management, and provide on-call troubleshooting and cross-team collaboration.
Top Skills:
AWSAzureContainersGCPGoGpusGrafanaKubernetesLinuxPrometheusPythonRayTpusVms
Information Technology
Build and maintain resilient infrastructure for the Intelligence Community. Responsibilities include implementing redundancy and monitoring, automating infrastructure and self-repair, reducing operational toil through scripting, improving security posture, and supporting cloud infrastructure. The role requires Linux systems engineering, software development, containerization, CI/CD, patching, system hardening, and a TS/SCI clearance with polygraph.
Top Skills:
Ccna-SecurityConfluenceDockerGitGoGsecJavaJenkinsJIRAKubernetesLinuxNessusNist 190Nist 800-53PackerPythonRhelRustSecurity+ CeSscp
Financial Services
Build and operate highly reliable clearing and risk systems supporting global financial markets. Responsibilities include architecting resilient infrastructure, automating lifecycle operations, embedding SRE practices into development, improving fault tolerance, leading observability and performance testing, preventing incidents, and guiding development and platform teams. The role requires cloud infrastructure expertise, coding proficiency, infrastructure as code, CI/CD, orchestration, configuration management, security, compliance, and strong cross-functional communication.
Top Skills:
BashChefCi/CdCloudFormationGkeGoGCPIaasJavaKubernetesOpentelemetryPaasPrometheusPythonRustTerraformTypescript
Fintech • Payments
Leads large-scale site reliability engineering strategy, architecting highly available and scalable systems, improving observability, automation, incident response, capacity planning, performance, and cloud costs. Builds self-healing mechanisms and AI agents that automate operational workflows, reduce TOIL, and support incident response and anomaly detection. Establishes AI security and governance controls, advises engineering leadership, leads cross-functional reliability initiatives, and mentors engineers developing production-grade SRE and agentic solutions.
Top Skills:
APIsCi/CdDistributed TracingDockerElk StackGrafanaJaegerKubernetesMySQLNoSQLOpentelemetryPostgresPrometheusService MeshesSplunk
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Operate and scale Kong’s global multi-region SaaS platform across AWS, GCP, and Azure. Build Kubernetes infrastructure, Terraform-based automation, Helm and ArgoCD deployment workflows, CI/CD pipelines, observability systems, and highly available data layers. Improve Kong Gateway and Mesh environments, reliability, scalability, security, disaster recovery, and cost efficiency. Participate in 24/7 on-call, incident response, SLO tracking, postmortems, and operational improvement initiatives.
Top Skills:
ArgocdAWSAzureAzure VnetBashCi/CdClickhouseDatadogDnsDruidGCPGcp NccGitopsGoGrafanaHelmHTTPKafkaKong GatewayKong MeshKubernetesLinuxLoad BalancersPostgresPrivatelinkPrometheusPythonRedisTerraformTerragruntThanosTls/SslTransit GatewayVpc Peering
Cloud
The Staff Site Reliability Engineer will manage large-scale cloud production systems, ensuring reliability and performance, while automating processes and responding to incidents.
Top Skills:
AWSBashCloudFormationDockerGoHelmKubernetesPythonRubyTerraform
Information Technology • Software • Travel
Drive reliability, autoscaling, and cost efficiency for cloud-based data and machine learning platforms. Manage Google Cloud services with Infrastructure as Code, optimize Kubernetes workloads, operationalize machine learning and Retrieval-Augmented Generation systems, and build CI automation, failovers, schema migrations, observability, and incident-response processes using Python and Bash.
Top Skills:
Apache AirflowBashBigQueryContinuous IntegrationGenerative AiGoogle Cloud PlatformGoogle Kubernetes EngineInfrastructure As CodePythonRetrieval-Augmented GenerationSQLTerraform
12 Days AgoSaved
Easy Apply
Easy Apply
Fintech • News + Entertainment • Software • Financial Services
Define tastytrade’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerts, observability standards, and production readiness reviews. Embed reliability patterns in Ruby, Java, and Elixir services running on HashiCorp Nomad. Extend Prometheus, Honeycomb, and OpenTelemetry observability; conduct fault-injection and tabletop exercises; strengthen on-call and incident-review processes; and mentor engineering teams in reliability practices.
Top Skills:
ConsulElixirGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Artificial Intelligence • Software
Architects and owns highly available infrastructure and Kubernetes-based platforms supporting autonomous systems. Builds Golang backend services, platform tooling, observability systems, dashboards, alerts, and log aggregation. Partners with product teams to launch services, performs performance analysis, manages cloud upgrades, and participates in incident response and postmortems. Collaborates on cloud security risk assessments, intrusion detection, threat-feed systems, risk mitigation, and SaaS payment processes. Provides architectural leadership and mentorship across engineering teams.
Top Skills:
ArgocdArgocd Image UpdaterArtifactoryAWSGithub ActionsGoJavaScriptKubernetesPythonRustTerraform
Fintech • Financial Services
Build and operate highly available, scalable, fault-tolerant production systems and observability platforms. Monitor system health, manage incidents, conduct blameless postmortems, improve testing and release procedures, support system design and capacity planning, and automate reliability improvements. Collaborate with engineering teams to establish service-level objectives and promote resilience engineering practices.
Top Skills:
CC++Cloud PlatformsDistributed SystemsGoJavaNetworkingObservability PlatformsPerlPythonRubyShell ScriptingUnix
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills:
AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results





.jpeg)




























