Top Site Reliability Engineer Jobs

Reposted 4 Days AgoSaved
In-Office
New York, NY, USA
120K-175K Annually
Senior level
120K-175K Annually
Senior level
Fintech • Financial Services
Design, build, and maintain reliable, scalable virtual desktop infrastructure (VDI) and supporting platforms. Lead incident response, automate deployments and operations with IaC and CI/CD, implement secure configurations, monitor system health, collaborate cross-functionally, and drive continuous improvement and operational excellence.
Top Skills: Active DirectoryAnsibleArm/BicepAzure DevopsCitrix CloudCitrix GatewayCvadDnsDscGithub ActionsGitlab CiGposJenkinsPowershellSsl/Tls CertificatesTerraformVdi Profile ManagementWindows 11 Multi-SessionWindows Server
4 Days AgoSaved
Remote
United States
125K-150K Annually
Mid level
125K-150K Annually
Mid level
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills: ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
4 Days AgoSaved
In-Office
Santa Barbara, CA, USA
175K-230K Annually
Expert/Leader
175K-230K Annually
Expert/Leader
Payments
Own platform reliability, security, automation, uptime, and incident response for payment-processing systems. Build and evolve AWS infrastructure with Terraform, manage infrastructure as code across cloud and physical environments, containerize legacy services, administer Linux servers, and embed PCI DSS compliance through patching, vulnerability management, logging, and monitoring. Provide technical leadership and support secure infrastructure for analytics and BI systems.
Top Skills: AnsibleAWSContainersLinuxOpenvoxPci DssPuppetTerraform
4 Days AgoSaved
Hybrid
5 Locations
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
Build and manage highly available Kubernetes platforms on AWS. Responsibilities include Kubernetes and Helm platform creation, Terraform infrastructure automation, Karpenter scaling, Istio service mesh management, CI/CD enablement, incident response, security, compliance, monitoring, cost optimization, troubleshooting, and documentation. The role supports multi-region cloud environments and requires federal-environment access eligibility, U.S. Person documentation, and occasional travel for in-person onboarding.
Top Skills: AnsibleAWSBashCi/CdCircleCICloudFormationCloudwatchDockerEc2EcsEksElk StackGitlabGoGrafanaHelmIamIstioJenkinsKarpenterKubernetesPrometheusPythonRdsS3SpinnakerTerraformVpc
5 Days AgoSaved
In-Office
2 Locations
106K-145K Annually
Senior level
106K-145K Annually
Senior level
Hardware • Information Technology • Other • Software • Analytics
Configure, optimize, and support AWS and Azure cloud infrastructure, especially AWS GovCloud environments supporting FedRAMP-compliant services. Improve availability, performance, security, scalability, and capacity; manage monitoring, application operations, incident response, runbooks, architecture documentation, and compliance. Collaborate with development teams and international stakeholders while using scripting, infrastructure-as-code, container, CI/CD, and observability tools.
Top Skills: AnsibleAWSAws GovcloudAzureBashEcsEksFedrampGitGrafanaJenkinsJIRAKubernetesPerlPowershellPrtgPythonSumo LogicTerraform
5 Days AgoSaved
In-Office
2 Locations
91K-125K Annually
Mid level
91K-125K Annually
Mid level
Hardware • Information Technology • Other • Software • Analytics
Design, build, and maintain scalable cloud infrastructure across AWS GovCloud and Azure. Develop automated deployment, monitoring, observability, incident response, and resilience testing systems. Establish SLOs, manage error budgets, troubleshoot application operations, and enforce FedRAMP security practices. Collaborate with architecture and software engineering teams to improve availability, performance, and automation across mission-critical SaaS and PaaS platforms.
Top Skills: Active DirectoryAnsibleAWSAws GovcloudAws SdkAzureBashDockerEcsEksFedrampGrafanaJira Service ManagementKubernetesNoSQLPerlPowershellPrtgPythonRdbmsSAMLSumoTerraform
Senior level
Financial Services
Lead application support and SRE activities for mission-critical systems, improving reliability, observability, resilience, and operational efficiency. Responsibilities include resolving production incidents, managing incident and problem processes, conducting root-cause reviews, supporting releases and disaster recovery testing, maintaining operational documentation, optimizing alerts, automating workflows, and strengthening risk and control practices. The role partners with engineering, infrastructure, operations, and global teams in a highly regulated financial services environment.
Top Skills: ActivemqAmazon Rds AuroraAutosysAws Ec2Aws IamAws LambdaAws S3Aws SqsBashCicsCobolDb2Db2 Stored ProceduresDynatraceFile-AidGrafanaIbm MqJavaScriptJclJIRAKafkaLinuxOpenshiftOracleOracle AqPerlPostgresPythonRabbitMQRubySeleniumServicenowShell ScriptingSnowflakeSplunkSpufiSQLWindows
5 Days AgoSaved
In-Office
3 Locations
90K-110K Annually
Senior level
90K-110K Annually
Senior level
Mining Operations
Leads reliability improvements for steel-producing equipment and facilities, reducing downtime through predictive and preventive maintenance programs. Uses Maximo and equipment databases, analyzes failure data, and supports criticality assessments, spare-parts analysis, root cause analysis, and FMEA. Provides troubleshooting and reliability expertise, collaborates on new and modified installations, updates engineering standards, and develops best practices through periodic travel to other company sites.
Top Skills: CmmsElectrical SystemsFmeaHydraulicsIbm Maximo EamMechanical SystemsPneumaticsReliability-Centered Maintenance
5 Days AgoSaved
Remote or Hybrid
USA
136K-181K Annually
Entry level
136K-181K Annually
Entry level
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills: Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
5 Days AgoSaved
In-Office
Alpharetta, GA, USA
Senior level
Senior level
Gaming • Mobile
Enhance production-system stability, performance, scalability, and reliability through observability, automation, Kubernetes operations, and AWS infrastructure management. Define and monitor SLIs, SLOs, and SLAs; manage incidents, on-call response, troubleshooting, root-cause analysis, and post-incident remediation. Improve CI/CD production readiness, secrets management, and application resilience while documenting runbooks and operational procedures.
Top Skills: .NetAmazon Ec2Amazon EksAmazon Route 53Amazon S3Argo CdAWSAws IamBashGithub ActionsGitlab Ci/CdGraylogHashicorp VaultHelmKubernetesNew RelicPackerPythonRancherTerraform
5 Days AgoSaved
Hybrid
Atlanta, GA, USA
115K-140K Annually
Senior level
115K-140K Annually
Senior level
Energy
Build and support scalable, resilient cloud-native platforms and applications. Responsibilities include reliability engineering, observability, incident response, platform engineering, cloud infrastructure automation, Infrastructure as Code, CI/CD enablement, disaster recovery, chaos engineering, and AI-driven operations. The role leads reliability improvements, develops self-service tooling, partners cross-functionally on platform modernization, and mentors junior engineers.
Top Skills: APIsAWSAzureBashChaos EngineeringCi/CdCloudFormationDatadogDistributed SystemsDockerGitGitopsGrafanaHelmInfrastructure As CodeKubernetesLinuxMicroservicesNew RelicOpentelemetryPrometheusPythonSplunkTerraform
Reposted 5 Days AgoSaved
In-Office
Orlando, FL, USA
175K-215K Annually
Senior level
175K-215K Annually
Senior level
Digital Media • Gaming • News + Entertainment • Sports
Lead multiple SRE teams to ensure reliability, scalability, and security across cloud and on-prem systems. Drive observability, automation, CI/CD, infrastructure-as-code, and operational excellence; set strategy, manage resources, mentor leaders, and influence stakeholders for commerce platforms.
Top Skills: Ai/MlAnsibleAWSAzureCi/CdCloudFormationGCPGitlabHarnessInfrastructure-As-CodeKubernetesServerlessTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
5 Days AgoSaved
In-Office
Barrington, RI, USA
Entry level
Entry level
Information Technology • Professional Services • Software • Consulting
Design, deploy, and maintain highly available production systems across cloud environments. Build automation and reliability tooling, manage Kubernetes and Docker workloads, implement Infrastructure as Code and CI/CD pipelines, and develop monitoring and observability using OpenTelemetry, Prometheus, and Grafana. Troubleshoot complex infrastructure and application issues, participate in incident response and root-cause analysis, improve system performance and capacity, and establish SRE reliability practices.
Top Skills: AnsibleAWSAzureCi/CdDockerGCPGoGrafanaJavaKubernetesLinuxOpentelemetryPrometheusPythonRustTerraformUnix
Reposted 10 Days AgoSaved
Hybrid
4 Locations
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills: AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Reposted 5 Days AgoSaved
Hybrid
San Francisco, CA, USA
200K-240K Annually
Senior level
200K-240K Annually
Senior level
Artificial Intelligence • Software
Design, build, and scale control- and data-plane infrastructure for distributed AI workloads. Improve reliability, performance, scheduling, and observability for Ray clusters across cloud and on-prem environments. Support accelerator integration, container image management, and provide on-call troubleshooting and cross-team collaboration.
Top Skills: AWSAzureContainersGCPGoGpusGrafanaKubernetesLinuxPrometheusPythonRayTpusVms
5 Days AgoSaved
In-Office
Aurora, CO, USA
87K-198K Annually
Senior level
87K-198K Annually
Senior level
Information Technology
Build and maintain resilient infrastructure for the Intelligence Community. Responsibilities include implementing redundancy and monitoring, automating infrastructure and self-repair, reducing operational toil through scripting, improving security posture, and supporting cloud infrastructure. The role requires Linux systems engineering, software development, containerization, CI/CD, patching, system hardening, and a TS/SCI clearance with polygraph.
Top Skills: Ccna-SecurityConfluenceDockerGitGoGsecJavaJenkinsJIRAKubernetesLinuxNessusNist 190Nist 800-53PackerPythonRhelRustSecurity+ CeSscp
5 Days AgoSaved
Hybrid
Wacker, IL, USA
132K-220K Annually
Expert/Leader
132K-220K Annually
Expert/Leader
Financial Services
Build and operate highly reliable clearing and risk systems supporting global financial markets. Responsibilities include architecting resilient infrastructure, automating lifecycle operations, embedding SRE practices into development, improving fault tolerance, leading observability and performance testing, preventing incidents, and guiding development and platform teams. The role requires cloud infrastructure expertise, coding proficiency, infrastructure as code, CI/CD, orchestration, configuration management, security, compliance, and strong cross-functional communication.
Top Skills: BashChefCi/CdCloudFormationGkeGoGCPIaasJavaKubernetesOpentelemetryPaasPrometheusPythonRustTerraformTypescript
6 Days AgoSaved
In-Office
3 Locations
121K-151K Annually
Senior level
121K-151K Annually
Senior level
Fintech • Payments
Leads large-scale site reliability engineering strategy, architecting highly available and scalable systems, improving observability, automation, incident response, capacity planning, performance, and cloud costs. Builds self-healing mechanisms and AI agents that automate operational workflows, reduce TOIL, and support incident response and anomaly detection. Establishes AI security and governance controls, advises engineering leadership, leads cross-functional reliability initiatives, and mentors engineers developing production-grade SRE and agentic solutions.
Top Skills: APIsCi/CdDistributed TracingDockerElk StackGrafanaJaegerKubernetesMySQLNoSQLOpentelemetryPostgresPrometheusService MeshesSplunk
6 Days AgoSaved
In-Office or Remote
Washington, DC, USA
123K-150K Annually
Entry level
123K-150K Annually
Entry level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Operate and scale Kong’s global multi-region SaaS platform across AWS, GCP, and Azure. Build Kubernetes infrastructure, Terraform-based automation, Helm and ArgoCD deployment workflows, CI/CD pipelines, observability systems, and highly available data layers. Improve Kong Gateway and Mesh environments, reliability, scalability, security, disaster recovery, and cost efficiency. Participate in 24/7 on-call, incident response, SLO tracking, postmortems, and operational improvement initiatives.
Top Skills: ArgocdAWSAzureAzure VnetBashCi/CdClickhouseDatadogDnsDruidGCPGcp NccGitopsGoGrafanaHelmHTTPKafkaKong GatewayKong MeshKubernetesLinuxLoad BalancersPostgresPrivatelinkPrometheusPythonRedisTerraformTerragruntThanosTls/SslTransit GatewayVpc Peering
Reposted 6 Days AgoSaved
Remote or Hybrid
3 Locations
174K-238K Annually
Senior level
174K-238K Annually
Senior level
Cloud
The Staff Site Reliability Engineer will manage large-scale cloud production systems, ensuring reliability and performance, while automating processes and responding to incidents.
Top Skills: AWSBashCloudFormationDockerGoHelmKubernetesPythonRubyTerraform
6 Days AgoSaved
Hybrid
Fort Worth, TX, USA
Mid level
Mid level
Information Technology • Software • Travel
Drive reliability, autoscaling, and cost efficiency for cloud-based data and machine learning platforms. Manage Google Cloud services with Infrastructure as Code, optimize Kubernetes workloads, operationalize machine learning and Retrieval-Augmented Generation systems, and build CI automation, failovers, schema migrations, observability, and incident-response processes using Python and Bash.
Top Skills: Apache AirflowBashBigQueryContinuous IntegrationGenerative AiGoogle Cloud PlatformGoogle Kubernetes EngineInfrastructure As CodePythonRetrieval-Augmented GenerationSQLTerraform
12 Days AgoSaved
Easy Apply
Hybrid
Chicago, IL, USA
Easy Apply
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Fintech • News + Entertainment • Software • Financial Services
Define tastytrade’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerts, observability standards, and production readiness reviews. Embed reliability patterns in Ruby, Java, and Elixir services running on HashiCorp Nomad. Extend Prometheus, Honeycomb, and OpenTelemetry observability; conduct fault-injection and tabletop exercises; strengthen on-call and incident-review processes; and mentor engineering teams in reliability practices.
Top Skills: ConsulElixirGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
7 Days AgoSaved
Remote
US
174K-305K Annually
Senior level
174K-305K Annually
Senior level
Artificial Intelligence • Software
Architects and owns highly available infrastructure and Kubernetes-based platforms supporting autonomous systems. Builds Golang backend services, platform tooling, observability systems, dashboards, alerts, and log aggregation. Partners with product teams to launch services, performs performance analysis, manages cloud upgrades, and participates in incident response and postmortems. Collaborates on cloud security risk assessments, intrusion detection, threat-feed systems, risk mitigation, and SaaS payment processes. Provides architectural leadership and mentorship across engineering teams.
Top Skills: ArgocdArgocd Image UpdaterArtifactoryAWSGithub ActionsGoJavaScriptKubernetesPythonRustTerraform
7 Days AgoSaved
In-Office
Dallas, TX, USA
Junior
Junior
Fintech • Financial Services
Build and operate highly available, scalable, fault-tolerant production systems and observability platforms. Monitor system health, manage incidents, conduct blameless postmortems, improve testing and release procedures, support system design and capacity planning, and automate reliability improvements. Collaborate with engineering teams to establish service-level objectives and promote resilience engineering practices.
Top Skills: CC++Cloud PlatformsDistributed SystemsGoJavaNetworkingObservability PlatformsPerlPythonRubyShell ScriptingUnix
Reposted 7 Days AgoSaved
In-Office or Remote
3 Locations
120K-261K Annually
Mid level
120K-261K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills: AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account