Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Reposted 28 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
As an intern, manage operational tasks in classified environments, develop automation tools, create documentation, and enhance services for Zscaler's cloud security platform.
Top Skills:
Aws EcsKubernetesPython
Digital Media • Information Technology • News + Entertainment
Responsible for ensuring reliability, scalability, and performance of data platforms. Design monitoring and alerting, automate deployments and recovery, optimize storage and query performance, troubleshoot incidents, plan capacity and scaling, document operations, enforce security/compliance, and collaborate with data engineering, product, and data science teams to maintain high availability of large-scale data systems.
Top Skills:
AnsibleAWSAzureCi/CdDockerElk StackGCPGoGrafanaJavaKubernetesMySQLNoSQLPostgresPrometheusPythonScalaTerraform
Digital Media • Information Technology • News + Entertainment
Responsible for ensuring reliability, scalability, and performance of data platforms: monitoring, incident response, automation, performance tuning, capacity planning, security/compliance, documentation, and cross-team collaboration to support large-scale data pipelines and backend data systems.
Top Skills:
AerospikeAnsibleAWSAws S3AzureCassandraCi/CdContainerizationDockerElk StackGCPGoGrafanaHadoopHdfsJavaKafkaKubernetesMicroservicesMySQLNoSQLPostgresPrometheusPythonScalaSnowflakeSparkTerraform
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Cloud • Information Technology • Internet of Things • Professional Services • Software
Lead design, build, and evolve developer infrastructure and CI platforms for Meraki cloud teams. Guide complex troubleshooting, mentor engineers, drive operational excellence, define roadmaps with leadership, and champion sustainable on-call practices while supporting large-scale distributed systems and automation across developer environments.
Top Skills:
Artifact ManagementBare MetalBuild ToolsCiCi PlatformsCode ReviewConfiguration-As-CodeContainer OrchestrationContainerizationInfrastructure AutomationManaged Cloud ServicesPythonRubyUnix/Linux
Retail
Ensures the reliability, availability, and performance of the 7NOW delivery platform through monitoring, incident response, automation, observability, and continuous improvement. Responsibilities include on-call support, root cause analysis, runbook development, cloud cost optimization, CI/CD maintenance, capacity planning, security remediation, disaster recovery testing, and performance tuning. The role also collaborates with engineering and infrastructure teams and mentors junior staff.
Top Skills:
AnsibleAWSAzureBashCi/CdCloudFormationDatadogDockerGithub ActionsGitlab CiGoGrafanaInfrastructure As CodeJenkinsKubernetesLinuxMicroservicesNew RelicNoSQLPowershellPrometheusPythonRest ApisSplunkSQLTerraformUnix
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills:
AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Cloud • Information Technology • Internet of Things • Professional Services • Software
Operate, deploy, and optimize backend collaboration services for global SaaS. Build and evolve CI/CD and automation, lead incident response and RCA, use observability for capacity planning, and define operational best practices and runbooks to improve reliability and scalability across cloud and hybrid environments.
Top Skills:
BashCi/CdDockerGitGoInfrastructure-As-CodeKubernetesLinuxMonitoringObservabilityPython
Cloud • Information Technology • Internet of Things • Professional Services • Software
Lead design, deployment, and operations of large-scale distributed databases (Cassandra, Kafka, OpenSearch) across AWS and private cloud. Automate infrastructure using Terraform, Ansible, and Kubernetes; implement CI/CD; perform capacity planning and performance tuning; manage backups, patching, and incident response for customer-facing SaaS; and consult with application teams on data access and cloud migration while participating in on-call rotations and ensuring compliance for US federal environments.
Top Skills:
AnsibleAWSCassandraGithub ActionsGitlab CiJenkinsKafkaKubernetesOpensearchPostgresPythonTerraform
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills:
Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Artificial Intelligence • Legal Tech • Software
As a Senior Site Reliability Engineer, you'll operate foundational platform services, enhance reliability standards, automate processes, and work with engineering teams to improve systems.
Top Skills:
Cloud InfrastructureKubernetesObservability Tools
Cloud • Software
Operate and improve large-scale on-premises infrastructure across bare-metal servers, VMs, storage, networking, Kubernetes, containers, CI/CD, and Kafka. Automate provisioning and configuration with Ansible and related infrastructure-as-code tools, maintain observability and reliability, troubleshoot hardware and operating systems, manage disaster recovery, and support incident response. The role requires onsite work weekly at a designated data center and participation in on-call rotations.
Top Skills:
AnsibleArgo CdBare-Metal ServersBashCertificate ManagementCi/CdContainersDhcpDnsElk StackFirewallsForemanGitopsGoGrafanaHelmKafkaKubernetesLinuxLinux NetworkingLoad BalancersMaasNtpPrometheusPythonRoutingStorage AppliancesTerraformVirtual MachinesVirtualizationVlans
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Fintech • Payments • Social Impact • Software
Build and manage backend applications and platform components for Chariot’s identity and payment network. Design scalable APIs and backend systems, establish engineering standards, improve monitoring and alerting, and collaborate directly with founders, product managers, and engineers. Drive projects from inception through delivery within an Agile/Scrum environment while contributing to architecture, technology strategy, and product decisions.
Top Skills:
AWSDockerGoGrpcKubernetesNode.jsPostgresRest ApisTerraform
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major incident response. Own remediation through closure, communicate incident updates, create playbooks and runbooks, and drive preventative improvements. Mentor SREs and develop automated operational processes for compliant financial platforms.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies, manage CI/CD pipelines, monitor CSPM alerts, and lead major-incident response as Incident Commander. Own remediation through closure, communicate incident status to technical and executive audiences, and develop incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and support compliant platforms meeting PCI-DSS and SOC2 standards.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer Terraform-based security policies, manage CI/CD pipelines, monitor CSPM alerts, and secure cloud workloads. Serve as Incident Commander during major incidents, coordinate technical responders, communicate status to stakeholders, track remediation through closure, and develop incident-management playbooks, runbooks, and escalation procedures. The role also involves building automated runbooks, mentoring SREs, and supporting PCI-DSS and SOC 2 compliance.
Top Skills:
AspmAWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWizZero-Trust Security
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer Terraform-based security baselines, optimize CI/CD pipelines, monitor cloud security posture, and lead major incident response as Incident Commander. Own remediation through closure, communicate incident updates to technical and executive audiences, and develop incident-management playbooks, runbooks, and escalation procedures. The role also involves architecting automated runbooks and mentoring SREs.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, lead major incidents, coordinate remediation through closure, communicate incident updates, and create incident-management playbooks and runbooks. The role also drives automated runbook architecture, mentors SREs, and supports compliance-focused platforms.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor cloud security posture with Wiz, and lead major-incident response as Incident Commander. Drive remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and develop automated runbooks for compliant, secure platforms.
Top Skills:
AWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWizZero-Trust Security
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools such as Wiz, and lead major incident response as Incident Commander. Own remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. The role also includes designing automated runbooks and mentoring SREs.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies, manage CI/CD pipelines, monitor cloud security posture, lead major incidents, coordinate remediation, communicate incident status, and develop incident-management playbooks and runbooks. The role also mentors SREs and supports compliant platforms using PCI-DSS and SOC 2 standards.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response. Own remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and drive complex automated runbook architecture.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoIamKubernetesPagerdutyPythonServicenowTerraformWizZero Trust
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response as Incident Commander. Own remediation through closure, communicate incident updates to technical and executive stakeholders, and develop incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and support compliant platforms subject to PCI-DSS and SOC 2 requirements.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer multi-cloud security policies, Terraform infrastructure, CI/CD pipelines, and automated runbooks. Monitor workloads with CSPM tools such as Wiz, lead major-incident response as Incident Commander, coordinate remediation through closure, communicate incident status, and author incident-management playbooks. Mentor SREs and support platforms subject to PCI-DSS and SOC2 compliance.
Top Skills:
AspmAWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPythonServicenowTerraformWizZero-Trust
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, manage CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response as Incident Commander. Own remediation through closure, communicate incident status to technical and executive audiences, and develop incident-management playbooks, runbooks, and escalation procedures. The role also involves designing automated runbooks, mentoring SREs, and supporting PCI-DSS and SOC2-compliant platforms.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results






















