Top Site Reliability Engineer Jobs

7 Days AgoSaved
In-Office
Dallas, TX, USA
Junior
Junior
Fintech • Financial Services
Build and operate highly available, scalable, fault-tolerant production systems and observability platforms. Monitor system health, manage incidents, conduct blameless postmortems, improve testing and release procedures, support system design and capacity planning, and automate reliability improvements. Collaborate with engineering teams to establish service-level objectives and promote resilience engineering practices.
Top Skills: CC++Cloud PlatformsDistributed SystemsGoJavaNetworkingObservability PlatformsPerlPythonRubyShell ScriptingUnix
Reposted 7 Days AgoSaved
In-Office or Remote
3 Locations
120K-261K Annually
Mid level
120K-261K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills: AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Reposted 12 Days AgoSaved
Easy Apply
Remote or Hybrid
9 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 7 Days AgoSaved
Hybrid
New York, NY, USA
95K-125K Annually
Mid level
95K-125K Annually
Mid level
Artificial Intelligence • eCommerce • Retail • Software
Build and maintain CI/CD pipelines, manage and automate cloud infrastructure and configurations, implement monitoring/logging and alerting for reliability, enforce security and compliance practices, and collaborate with development teams to support scaling and operations.
Top Skills: Soc ISoc Ii
7 Days AgoSaved
Hybrid
New York, NY, USA
200K-250K Annually
Expert/Leader
200K-250K Annually
Expert/Leader
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Leads reliability engineering for a cloud-based, real-time healthcare platform. Responsibilities include maintaining 99.99% uptime, designing fault-tolerant and self-healing systems, managing AWS and Kubernetes infrastructure, implementing Terraform and observability tooling, leading incident response and disaster recovery, ensuring HIPAA and SOC 2 compliance, and driving scalability strategy while mentoring the SRE team.
Top Skills: AWSCi/CdDatadogDistributed SystemsDockerGitopsGrafanaKubernetesOpentelemetryPrometheusTerraform
Reposted 7 Days AgoSaved
In-Office
2 Locations
150K-225K Annually
Mid level
150K-225K Annually
Mid level
Cloud • Information Technology • Security • Software
Provide 24/7 technical support for a SaaS AI security platform, monitor uptime, triage and resolve incidents, collaborate with customers and development teams, and drive SRE-focused automation and reliability improvements.
Top Skills: Ai InferenceAWSGrafanaHTTPJSONKubernetesLinuxPostgresPrometheusPythonRest ApisSaaSTerraformTicketing System
Reposted 7 Days AgoSaved
Hybrid
5 Locations
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills: AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
Reposted 7 Days AgoSaved
Hybrid
Denver, CO, USA
190K-215K Annually
Expert/Leader
190K-215K Annually
Expert/Leader
Artificial Intelligence • Information Technology • Software • Infrastructure as a Service (IaaS)
Own fleet-scale reliability, upgradeability, and operational excellence for an edge platform. Design and operate automated, secure lifecycle systems, observability, and incident response. Lead cross-domain, high-severity incident ownership, set production standards (SLOs/SLIs, change management), mentor engineers, and apply AI to accelerate diagnostics and operational workflows.
Top Skills: Ai SystemsAudit LoggingCanary DeploymentsConfiguration ManagementFleet ManagementInfrastructure-As-CodeKubernetesLinuxObservabilitySecure BootStaged Rollouts
7 Days AgoSaved
In-Office
New Albany, IN, USA
100K-150K Annually
Senior level
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Leads site reliability engineering for critical production systems, focusing on operational excellence, automation, performance optimization, scalability, observability, and technical incident response. Builds internal tooling and self-healing systems, manages cloud and Kubernetes-based environments, analyzes monitoring and system metrics, and drives data-backed reliability and capacity decisions while collaborating with technical and executive stakeholders.
Top Skills: AWSAzureDynatraceElk StackGCPKubernetesMicroservicesPythonSplunk
7 Days AgoSaved
Hybrid
2 Locations
Senior level
Senior level
Fintech • Financial Services
Design and implement observability across distributed systems, integrating monitoring tools and building dashboards, alerts, and telemetry frameworks. Automate repetitive operational work through self-healing and auto-remediation solutions. Support production incident triage, troubleshooting, and root cause analysis while improving reliability, performance, scalability, and capacity planning. Collaborate with engineering and support teams to strengthen system stability and operational efficiency.
Top Skills: AppdynamicsBashCi/CdCloud PlatformsDevOpsDistributed SystemsDynatraceGrafanaJavaKubernetesMicroservicesPythonSlisSlosSplunk
Reposted 7 Days AgoSaved
Hybrid
Mountain View, CA, USA
252K-308K Annually
Senior level
252K-308K Annually
Senior level
Fintech • Payments • Financial Services
Lead EarnIn's AI-first reliability engineering, enhancing incident response, automation, and resilience in operations while mentoring engineers.
Top Skills: AIAWSCloudwatchDatadogGoKubernetesOpentelemetryPythonTerraform
Reposted 7 Days AgoSaved
In-Office or Remote
9 Locations
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 7 Days AgoSaved
In-Office
5 Locations
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills: AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Reposted 7 Days AgoSaved
In-Office
Frederick, MD, USA
140K-155K Annually
Senior level
140K-155K Annually
Senior level
Information Technology • Consulting
The Site Reliability Engineer designs monitoring frameworks, manages SLIs/SLOs, automates cloud infrastructure, ensures compliance, and promotes SRE best practices.
Top Skills: .Net CoreAmazon AwsAnsibleBashChefDockerElkGithub ActionsGoogle GcpGrafanaJavaJenkinsKubernetesAzureNode.jsOpen-TelemetryPowershellPrometheusPuppetPythonRSplunkTerraform
8 Days AgoSaved
In-Office or Remote
2 Locations
192K-356K Annually
Expert/Leader
192K-356K Annually
Expert/Leader
Cloud • Information Technology • Internet of Things • Professional Services • Software
Leads technical strategy for site reliability, critical incident response, and production resilience across Splunk Cloud. Serves as the escalation authority during P1/P2 events, owns complex enterprise customer environments, guides infrastructure architecture and automation, leads post-mortems and root-cause analysis, improves observability and operational processes, and mentors senior engineers without formal management authority.
Top Skills: AWSGoGoogle Cloud Platform (Gcp)Indexer ClusteringKvstoreLinuxAzurePythonSearch Head ClustersSplSplunkSplunk Observability
8 Days AgoSaved
Hybrid
New York, NY, USA
200K-250K Annually
Expert/Leader
200K-250K Annually
Expert/Leader
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Lead reliability, performance, automation, and scalability for a 24/7 AWS-based healthcare platform. Responsibilities include designing fault-tolerant infrastructure, managing Kubernetes and Docker environments, implementing Terraform and GitOps, improving observability, leading incident response and disaster recovery, and ensuring HIPAA and SOC 2 compliance. The role also sets technical strategy, reduces downtime through self-healing automation, and leads and mentors the SRE team.
Top Skills: AWSCi/CdDatadogDockerFhirGitopsGrafanaHipaaHl7KubernetesOpentelemetryPrometheusSoc 2Terraform
8 Days AgoSaved
In-Office
North Tower Forest, TN, USA
Senior level
Senior level
Financial Services
Support and mature the cybersecurity Site Reliability Engineering capability for critical security platforms. Operate SIEM, XDR, CyberArk, SOAR, NDR, IPS, and related systems; perform monitoring, incident response, problem management, and root cause analysis. Build automation with Python, Ansible, APIs, and infrastructure-as-code. Maintain observability dashboards, define SLIs and SLOs, document runbooks, support production handovers, participate in on-call duties, and conduct disaster recovery and resilience testing.
Top Skills: AnsibleAPIsCyberarkGrafanaHipsInfrastructure As CodeIpsLinuxNdrPowershellPrometheusPythonSIEMSoarSplunkSyslogTipWindowsXdr
Reposted 8 Days AgoSaved
In-Office
Chicago, IL, USA
Senior level
Senior level
Information Technology • Professional Services • Consulting
Design, implement, and maintain scalable AWS infrastructure and CI/CD pipelines using Terraform and Harness. Build observability with Dynatrace/Datadog, troubleshoot production incidents, run blameless postmortems, optimize cost/performance, drive chaos engineering, embed SRE practices (SLAs/SLOs), and mentor junior engineers.
Top Skills: AnsibleApp MeshAWSAws Well-Architected FrameworkCi/CdCloudfrontContainerizationDatadogDynatraceEc2EcsEksGoHarnessIamInfrastructure As CodeIstioKubernetesLambdaLinuxPythonRdsS3ShellTerraformVpc
Reposted 8 Days AgoSaved
In-Office
San Francisco, CA, USA
116K-200K Annually
Mid level
116K-200K Annually
Mid level
Information Technology • Mobile • Software
As a Site Reliability Engineer, you'll ensure system reliability and scalability, automate processes, optimize performance, and collaborate on system design.
Top Skills: AWSAzureBashCloudFormationDatadogDockerElkGoGoogle Cloud PlatformGrafanaHelmKubernetesNew RelicPrometheusPulumiPythonTerraform
Reposted 8 Days AgoSaved
In-Office
San Francisco, CA, USA
230K-310K Annually
Senior level
230K-310K Annually
Senior level
Artificial Intelligence • Software
Own the reliability and performance of backend systems at Gamma, building automation and tooling while leading incident response and improving system stability.
Top Skills: AWSCloudFormationDockerGoKafkaKubernetesNode.jsPythonTerraformTypescript
Reposted 8 Days AgoSaved
In-Office
San Francisco, CA, USA
Mid level
Mid level
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own reliability for a live fleet of Linux-based edge sensors and cloud infrastructure. Triage and recover field hardware, perform SSH-based diagnostics, build fleet management and OTA systems, implement observability and alerting, automate operational tasks, develop runbooks, and participate in on-call rotations to prevent and resolve incidents.
Top Skills: AWSBashCDnsDockerFirewallsGoIamKubernetesLinuxPythonRustSshVpn
Reposted 8 Days AgoSaved
In-Office
3 Locations
72K-119K Annually
Junior
72K-119K Annually
Junior
Information Technology • Legal Tech • Analytics
Support and automate cloud infrastructure (primarily Azure) for insurance applications. Implement IaC, assist deployments, troubleshoot incidents, improve security controls, and maintain monitoring and documentation while collaborating with engineers and non-technical stakeholders.
Top Skills: Arm TemplatesAWSAzureBashCi/CdContainerizationDnsGCPGitGitMonitoring And Logging ToolsPowershellPythonServerless FunctionsTerraformVnets
Reposted 8 Days AgoSaved
In-Office or Remote
San Mateo, CA, USA
140K-210K Annually
Junior
140K-210K Annually
Junior
Artificial Intelligence • Hardware • Robotics • Software
As a Software Engineer - Infrastructure at Skydio, you'll support and enhance Kubernetes infrastructure, improve continuous delivery, and collaborate on security and architectural improvements.
Top Skills: Cloud PlatformsContinuous DeploymentGoInfrastructureKubernetesPython
Reposted 8 Days AgoSaved
In-Office
San Jose, CA, USA
149K-361K Annually
Senior level
149K-361K Annually
Senior level
News + Entertainment
As a Senior Machine Learning Engineer, you will develop advanced machine learning and deep learning models and platforms for optimizing advertising performance and conduct complex experiments.
Top Skills: AIControl SystemsDeep LearningMachine LearningReinforcement LearningStatistical Techniques
Reposted 8 Days AgoSaved
In-Office
Austin, TX, USA
Senior level
Senior level
News + Entertainment
Design, operate, and scale cloud-native ML infrastructure across GCP and AWS (GPU/TPU), build CI/CD for models, maintain low-latency real-time inference systems, define observability and monitoring for ML models, participate in on-call incident response, and partner with data scientists to improve MLOps and platform usability.
Top Skills: AerospikeApache AirflowApache FlinkSparkAWSChrononDatadogEksGCPGitlab RunnerGkeGpuGrafanaJavaJenkinsKafkaKubernetesKv StoreMlflowPrometheusPythonRayScalaTerraformTpuVector Database
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account