Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies, manage CI/CD pipelines, monitor cloud security posture, lead major incidents, coordinate remediation, communicate incident status, and develop incident-management playbooks and runbooks. The role also mentors SREs and supports compliant platforms using PCI-DSS and SOC 2 standards.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response. Own remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and drive complex automated runbook architecture.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoIamKubernetesPagerdutyPythonServicenowTerraformWizZero Trust
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response as Incident Commander. Own remediation through closure, communicate incident updates to technical and executive stakeholders, and develop incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and support compliant platforms subject to PCI-DSS and SOC 2 requirements.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer multi-cloud security policies, Terraform infrastructure, CI/CD pipelines, and automated runbooks. Monitor workloads with CSPM tools such as Wiz, lead major-incident response as Incident Commander, coordinate remediation through closure, communicate incident status, and author incident-management playbooks. Mentor SREs and support platforms subject to PCI-DSS and SOC2 compliance.
Top Skills:
AspmAWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPythonServicenowTerraformWizZero-Trust
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, manage CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response as Incident Commander. Own remediation through closure, communicate incident status to technical and executive audiences, and develop incident-management playbooks, runbooks, and escalation procedures. The role also involves designing automated runbooks, mentoring SREs, and supporting PCI-DSS and SOC2-compliant platforms.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies, manage CI/CD pipelines, monitor workloads with CSPM tools such as Wiz, and lead major incident response as Incident Commander. Own remediation through closure, communicate with technical and executive stakeholders, and develop incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and support compliant platforms subject to PCI-DSS and SOC 2 requirements.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build secure infrastructure with Terraform, optimize CI/CD pipelines, monitor workloads using CSPM tools such as Wiz, and lead major incident response. Own remediation through closure, communicate incident updates to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. The role also drives automated runbook architecture and mentors SREs.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintains operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineers Terraform-based security policies, manages CI/CD pipelines, monitors CSPM alerts, and leads major-incident response as Incident Commander. Owns remediation through closure, communicates incident updates to technical and executive audiences, and develops incident-management playbooks, runbooks, and escalation procedures. The role also mentors SREs and supports compliant platforms subject to PCI-DSS and SOC2 standards.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies and configuration baselines, optimize CI/CD pipelines, monitor cloud security posture, and secure workloads with CSPM tools. Serve as Incident Commander for major incidents, coordinate responders, communicate status, track remediation through closure, and develop incident-management playbooks and runbooks. The role also drives automated runbook architecture and mentors SREs while supporting compliance requirements.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build multi-cloud security baselines with Terraform, optimize CI/CD pipelines, monitor workloads using CSPM tools such as Wiz, and lead major incident response as Incident Commander. Own remediation through closure, communicate incident updates, develop playbooks and runbooks, and drive preventative improvements while mentoring SREs.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies, configuration baselines, CI/CD pipelines, and automated runbooks. Monitor cloud security alerts using CSPM tools such as Wiz, lead major-incident response as Incident Commander, coordinate remediation through closure, communicate with technical and executive stakeholders, and develop incident-management playbooks and escalation procedures.
Top Skills:
AspmAWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPythonServicenowTerraformWizZero Trust
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor cloud security posture, and lead major incident response. Own remediation through closure, communicate with technical and executive stakeholders, and create incident-management playbooks and runbooks. The role also drives automated runbook architecture, mentors SREs, and supports compliance-focused platforms.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPythonServicenowTerraformWiz
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with Wiz, lead major incidents, coordinate remediation, communicate status to stakeholders, and create incident-management playbooks and runbooks. The role also mentors SREs and supports platforms subject to PCI-DSS and SOC 2 compliance.
Top Skills:
AWSAzureCi/CdCnappCspmGoGoogle Cloud PlatformKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major incidents as Incident Commander. Own remediation through closure, communicate incident status to technical and executive audiences, and develop incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and support compliant platforms subject to PCI-DSS and SOC 2 standards.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies and configuration baselines, optimize CI/CD pipelines, monitor cloud security posture, and lead major incident response. Own remediation through closure, communicate incident updates to technical and executive stakeholders, and create incident-management playbooks and runbooks. The role also architects automated response processes, mentors SREs, and supports compliance with PCI-DSS and SOC 2 standards.
Top Skills:
AspmAWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security policies, manage CI/CD pipelines, monitor workloads with CSPM tools such as Wiz, and lead major incidents as Incident Commander. Own remediation through closure, communicate incident status to technical and executive audiences, and create incident-management playbooks, runbooks, and escalation procedures. The role also includes designing automated runbooks and mentoring SREs.
Top Skills:
AWSAzureCi/CdCnappCspmFederated IamGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWizZero Trust
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead architecture, modernization, and reliability of mainframe CICS/MQ and z/OS Connect platforms. Provide technical leadership, performance tuning, incident resolution, governance, roadmaps, and collaborate with stakeholders to integrate and modernize mainframe workloads.
Top Skills:
CicsCobolIbm MqIbm Z/OsZ/Os Connect
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Lead architecture, modernization, and reliability of mainframe z/OS Db2 systems. Provide technical leadership, drive automation and CI/CD for database changes, monitor performance using SMF/RMF, troubleshoot incidents, and collaborate with stakeholders to ensure high availability, security, and scalability while guiding modernization efforts.
Top Skills:
AnsibleCi/CdDb2Ibm Z/OsOpenshiftPythonRed Hat Ansible For Ibm Z CollectionsRmfSmfSql (Ddl/Dml)Zlinux
Artificial Intelligence • Transportation
Build and improve reliability tooling, automation, observability, and incident-management processes for autonomous vehicle software. Diagnose complex production issues across Linux, systems, logs, metrics, traces, and vehicle data; develop solutions that accelerate detection, triage, and recovery. Collaborate with software, field engineering, and operations teams while supporting safety-operator tools and partner vehicle programs. The role requires production coding and hands-on systems debugging in a hybrid Sunnyvale environment.
Top Skills:
C++Ci/CdContainerizationDatabasesDatadogDistributed SystemsGrafanaHumioLinuxNetworkingOpentelemetryPrometheusPythonRustSplunk
Fintech • Financial Services
Lead Site Reliability Engineer responsible for production support, automating deployments, monitoring availability and performance, troubleshooting infrastructure and applications, driving reliability improvements, collaborating with development and infrastructure teams, and participating in 24/7 on-call rotation.
Top Skills:
AutosysAWSAzureC#Ci/CdContainersDb2Generative Ai ToolsIp SoftJavaJenkinsLinuxMqOraclePerlPythonRubyShellSockeyeSplunkSybaseTrainUnixVirtual MachinesWeb ServicesWindows
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills:
AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
Big Data • Cloud • Software • Database
The Site Reliability Engineer designs and builds infrastructure for a global cloud service, implements automation, and optimizes system performance while managing on-call operations.
Top Skills:
AWSDnsGCPHTTPKubernetesLinuxAzureProgramming LanguagesTls
Fintech • Financial Services
Lead and improve production reliability for Wealth Management services: monitor availability and performance, automate deployments and operational tasks, troubleshoot infrastructure and databases, collaborate with development and upstream teams, maintain runbooks/knowledgebase, perform root cause analysis, and participate in 24/7 on-call coverage and offshore coordination to reduce outages and operational toil.
Top Skills:
AutosysAWSAzureC#Ci/CdContainersDb2Ip SoftJavaJenkinsLinuxMqOraclePerlPythonRubyShellSockeyeSplunkSybaseTrainUnixVirtual MachinesWeb ServicesWindeployWindows
Cloud • Information Technology • Consulting • Design • Generative AI
Build and operate AWS infrastructure, including EKS, networking, security, hybrid connectivity, and monitoring. Automate infrastructure and deployments with Terraform, GitLab CI/CD, shell scripting, and Python. Manage Kubernetes and Istio, observability with CloudWatch, Grafana, and OpenTelemetry, plus patching and vulnerability remediation. Support release engineering and monitoring workflows using Kibana, Dynatrace, and Apigee in an onsite Agile/Scrum environment.
Top Skills:
ApigeeAWSCloudwatchDirect ConnectDynatraceEc2EksGitlab Ci/CdGrafanaGuarddutyIamIstioKibanaKubernetesMskOpentelemetryPythonRoute 53S3Security HubShellSsoTerraformTransit GatewayVpc
Logistics • Mobile • Productivity • Software • Transportation
The Senior Site Reliability Engineer will manage the reliability of Zello's data tier, contribute to monitoring and incident response while improving cloud infrastructure and database performance.
Top Skills:
BashDockerElasticsearchGoKubernetesLokiMongoDBMySQLPrometheusPythonRedisScylladbTempo
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results


















