Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Fintech • Financial Services
Build and operate highly available, scalable, fault-tolerant production systems and observability platforms. Monitor system health, manage incidents, conduct blameless postmortems, improve testing and release procedures, support system design and capacity planning, and automate reliability improvements. Collaborate with engineering teams to establish service-level objectives and promote resilience engineering practices.
Top Skills:
CC++Cloud PlatformsDistributed SystemsGoJavaNetworkingObservability PlatformsPerlPythonRubyShell ScriptingUnix
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills:
AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Reposted 12 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Artificial Intelligence • eCommerce • Retail • Software
Build and maintain CI/CD pipelines, manage and automate cloud infrastructure and configurations, implement monitoring/logging and alerting for reliability, enforce security and compliance practices, and collaborate with development teams to support scaling and operations.
Top Skills:
Soc ISoc Ii
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Leads reliability engineering for a cloud-based, real-time healthcare platform. Responsibilities include maintaining 99.99% uptime, designing fault-tolerant and self-healing systems, managing AWS and Kubernetes infrastructure, implementing Terraform and observability tooling, leading incident response and disaster recovery, ensuring HIPAA and SOC 2 compliance, and driving scalability strategy while mentoring the SRE team.
Top Skills:
AWSCi/CdDatadogDistributed SystemsDockerGitopsGrafanaKubernetesOpentelemetryPrometheusTerraform
Cloud • Information Technology • Security • Software
Provide 24/7 technical support for a SaaS AI security platform, monitor uptime, triage and resolve incidents, collaborate with customers and development teams, and drive SRE-focused automation and reliability improvements.
Top Skills:
Ai InferenceAWSGrafanaHTTPJSONKubernetesLinuxPostgresPrometheusPythonRest ApisSaaSTerraformTicketing System
Cloud
Design, automate, and maintain highly available cloud infrastructure and CI/CD platforms. Troubleshoot and debug Linux systems and networking, automate infrastructure (Terraform/Chef/Ansible/Puppet), and improve scalability, reliability, and platform velocity. Participate in on-call rotation and collaborate across engineering teams.
Top Skills:
AnsibleApache HttpdApache TomcatAWSBashChefCi/CdDnsDockerFedrampGdbGitGoHTTPKubernetesLinuxLoad BalancingLtraceNginxPki/Federated Certificate ManagementPuppetPythonStraceTcp/IpTcpdumpTerraformWireshark
Artificial Intelligence • Information Technology • Software • Infrastructure as a Service (IaaS)
Own fleet-scale reliability, upgradeability, and operational excellence for an edge platform. Design and operate automated, secure lifecycle systems, observability, and incident response. Lead cross-domain, high-severity incident ownership, set production standards (SLOs/SLIs, change management), mentor engineers, and apply AI to accelerate diagnostics and operational workflows.
Top Skills:
Ai SystemsAudit LoggingCanary DeploymentsConfiguration ManagementFleet ManagementInfrastructure-As-CodeKubernetesLinuxObservabilitySecure BootStaged Rollouts
Artificial Intelligence • Information Technology • Software • Consulting
Leads site reliability engineering for critical production systems, focusing on operational excellence, automation, performance optimization, scalability, observability, and technical incident response. Builds internal tooling and self-healing systems, manages cloud and Kubernetes-based environments, analyzes monitoring and system metrics, and drives data-backed reliability and capacity decisions while collaborating with technical and executive stakeholders.
Top Skills:
AWSAzureDynatraceElk StackGCPKubernetesMicroservicesPythonSplunk
Fintech • Financial Services
Design and implement observability across distributed systems, integrating monitoring tools and building dashboards, alerts, and telemetry frameworks. Automate repetitive operational work through self-healing and auto-remediation solutions. Support production incident triage, troubleshooting, and root cause analysis while improving reliability, performance, scalability, and capacity planning. Collaborate with engineering and support teams to strengthen system stability and operational efficiency.
Top Skills:
AppdynamicsBashCi/CdCloud PlatformsDevOpsDistributed SystemsDynatraceGrafanaJavaKubernetesMicroservicesPythonSlisSlosSplunk
Fintech • Payments • Financial Services
Lead EarnIn's AI-first reliability engineering, enhancing incident response, automation, and resilience in operations while mentoring engineers.
Top Skills:
AIAWSCloudwatchDatadogGoKubernetesOpentelemetryPythonTerraform
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills:
AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills:
AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Information Technology • Consulting
The Site Reliability Engineer designs monitoring frameworks, manages SLIs/SLOs, automates cloud infrastructure, ensures compliance, and promotes SRE best practices.
Top Skills:
.Net CoreAmazon AwsAnsibleBashChefDockerElkGithub ActionsGoogle GcpGrafanaJavaJenkinsKubernetesAzureNode.jsOpen-TelemetryPowershellPrometheusPuppetPythonRSplunkTerraform
Cloud • Information Technology • Internet of Things • Professional Services • Software
Leads technical strategy for site reliability, critical incident response, and production resilience across Splunk Cloud. Serves as the escalation authority during P1/P2 events, owns complex enterprise customer environments, guides infrastructure architecture and automation, leads post-mortems and root-cause analysis, improves observability and operational processes, and mentors senior engineers without formal management authority.
Top Skills:
AWSGoGoogle Cloud Platform (Gcp)Indexer ClusteringKvstoreLinuxAzurePythonSearch Head ClustersSplSplunkSplunk Observability
Cloud • Healthtech • Internet of Things • Machine Learning • Software
Lead reliability, performance, automation, and scalability for a 24/7 AWS-based healthcare platform. Responsibilities include designing fault-tolerant infrastructure, managing Kubernetes and Docker environments, implementing Terraform and GitOps, improving observability, leading incident response and disaster recovery, and ensuring HIPAA and SOC 2 compliance. The role also sets technical strategy, reduces downtime through self-healing automation, and leads and mentors the SRE team.
Top Skills:
AWSCi/CdDatadogDockerFhirGitopsGrafanaHipaaHl7KubernetesOpentelemetryPrometheusSoc 2Terraform
Financial Services
Support and mature the cybersecurity Site Reliability Engineering capability for critical security platforms. Operate SIEM, XDR, CyberArk, SOAR, NDR, IPS, and related systems; perform monitoring, incident response, problem management, and root cause analysis. Build automation with Python, Ansible, APIs, and infrastructure-as-code. Maintain observability dashboards, define SLIs and SLOs, document runbooks, support production handovers, participate in on-call duties, and conduct disaster recovery and resilience testing.
Top Skills:
AnsibleAPIsCyberarkGrafanaHipsInfrastructure As CodeIpsLinuxNdrPowershellPrometheusPythonSIEMSoarSplunkSyslogTipWindowsXdr
Information Technology • Professional Services • Consulting
Design, implement, and maintain scalable AWS infrastructure and CI/CD pipelines using Terraform and Harness. Build observability with Dynatrace/Datadog, troubleshoot production incidents, run blameless postmortems, optimize cost/performance, drive chaos engineering, embed SRE practices (SLAs/SLOs), and mentor junior engineers.
Top Skills:
AnsibleApp MeshAWSAws Well-Architected FrameworkCi/CdCloudfrontContainerizationDatadogDynatraceEc2EcsEksGoHarnessIamInfrastructure As CodeIstioKubernetesLambdaLinuxPythonRdsS3ShellTerraformVpc
Information Technology • Mobile • Software
As a Site Reliability Engineer, you'll ensure system reliability and scalability, automate processes, optimize performance, and collaborate on system design.
Top Skills:
AWSAzureBashCloudFormationDatadogDockerElkGoGoogle Cloud PlatformGrafanaHelmKubernetesNew RelicPrometheusPulumiPythonTerraform
Artificial Intelligence • Software
Own the reliability and performance of backend systems at Gamma, building automation and tooling while leading incident response and improving system stability.
Top Skills:
AWSCloudFormationDockerGoKafkaKubernetesNode.jsPythonTerraformTypescript
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own reliability for a live fleet of Linux-based edge sensors and cloud infrastructure. Triage and recover field hardware, perform SSH-based diagnostics, build fleet management and OTA systems, implement observability and alerting, automate operational tasks, develop runbooks, and participate in on-call rotations to prevent and resolve incidents.
Top Skills:
AWSBashCDnsDockerFirewallsGoIamKubernetesLinuxPythonRustSshVpn
Information Technology • Legal Tech • Analytics
Support and automate cloud infrastructure (primarily Azure) for insurance applications. Implement IaC, assist deployments, troubleshoot incidents, improve security controls, and maintain monitoring and documentation while collaborating with engineers and non-technical stakeholders.
Top Skills:
Arm TemplatesAWSAzureBashCi/CdContainerizationDnsGCPGitGitMonitoring And Logging ToolsPowershellPythonServerless FunctionsTerraformVnets
Artificial Intelligence • Hardware • Robotics • Software
As a Software Engineer - Infrastructure at Skydio, you'll support and enhance Kubernetes infrastructure, improve continuous delivery, and collaborate on security and architectural improvements.
Top Skills:
Cloud PlatformsContinuous DeploymentGoInfrastructureKubernetesPython
News + Entertainment
As a Senior Machine Learning Engineer, you will develop advanced machine learning and deep learning models and platforms for optimizing advertising performance and conduct complex experiments.
Top Skills:
AIControl SystemsDeep LearningMachine LearningReinforcement LearningStatistical Techniques
News + Entertainment
Design, operate, and scale cloud-native ML infrastructure across GCP and AWS (GPU/TPU), build CI/CD for models, maintain low-latency real-time inference systems, define observability and monitoring for ML models, participate in on-call incident response, and partner with data scientists to improve MLOps and platform usability.
Top Skills:
AerospikeApache AirflowApache FlinkSparkAWSChrononDatadogEksGCPGitlab RunnerGkeGpuGrafanaJavaJenkinsKafkaKubernetesKv StoreMlflowPrometheusPythonRayScalaTerraformTpuVector Database
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results

































