Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Automotive • Information Technology • Logistics • Software
Lead Site Reliability Engineer implements IaC and automation, builds observability (SLIs/SLOs, dashboards, alerting), manages incident response, runbooks, gamedays, postmortems, and drives SRE/DevOps best practices, AppSec integration, testing, and CI/CD improvements across teams.
Top Skills:
AppsecAWSAws CloudformationC#Ci/CdCloudsploitCloudwatchData TheoremDatadogGrafanaIacInfrastructure As CodeJavaNewrelicPythonTerraformVeracode
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills:
AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Edtech • Information Technology • Software
Lead infrastructure, reliability, and observability across multi-cloud environments. Improve CI/CD, IaC standards, staging parity, Kubernetes operations, monitoring and SLOs, incident response, and platform modernization while partnering with engineering teams.
Top Skills:
Ai-Assisted Development Tools (Claude CodeAutoscalingAWSCi/Cd PipelinesCodex)Event-Driven ArchitecturesGCPIncident ManagementInfrastructure-As-CodeKubernetesMonitoringObservabilityPythonQueue-Based ArchitecturesRuby On RailsSlo FrameworksTerraform
Artificial Intelligence • Software • Generative AI
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating cloud-native systems, owning production reliability and incident response, building observability and automation, defining SLOs/SLIs, driving postmortems, and partnering with product and engineering teams to improve operational maturity.
Top Skills:
AWSAzureGCPGoJavaKubernetesPython
Security • Software • Cybersecurity
Hands-on Site Reliability Engineer responsible for building and maintaining cloud infrastructure, CI/CD pipelines, observability (logging/monitoring/tracing), automation, and security best practices. Manage datacenter resources, troubleshoot clusters and services, collaborate with engineering teams for deployments, and participate in on-call incident response to ensure high availability and performance.
Top Skills:
AnsibleArgocdBashChefDatadogElkGitlab CiGoGrafanaJenkinsKubernetesLinuxPrometheusPythonRancher
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills:
Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills:
Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Automotive • eCommerce • Retail • Sales
Lead SRE enablement by defining SLO/SLO frameworks, production readiness, and reliability playbooks. Build and standardize observability (Dynatrace), provide alerting/dashboard/runbook templates, coach teams on SRE practices, run training, participate in incident post-mortems, and report enterprise reliability metrics while advising on architecture for hybrid GCP and on-prem environments.
Top Skills:
AnsibleApmDynatraceGoGoogle Cloud Platform (Gcp)JavaKubernetesObservabilityPythonTerraform
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills:
AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
eCommerce • Fashion
Leads enterprise observability and Site Reliability Engineering strategy, overseeing real-time monitoring, reliability practices, automation, self-healing, performance optimization, and operational excellence. Builds and develops a global team while partnering with engineering, infrastructure, cybersecurity, architecture, product, and business leaders. Establishes reliability standards, drives predictive service insights, advises executives on technology risk and service health, and delivers measurable improvements in availability, scalability, performance, and incident resolution.
Top Skills:
Anomaly DetectionAutomationChaos TestingCloud OperationsDashboardsInfrastructureItilObservabilityService ManagementSite Reliability Engineering (Sre)
Big Data • Real Estate • Software
Senior SRE responsible for reliability, observability, and operational excellence of a large AWS/Kubernetes platform. Duties include maintaining EKS/Fargate infrastructure, monitoring SLIs/SLOs, implementing observability with NewRelic, driving cost optimization and FinOps practices, executing chaos engineering and incident response, contributing automation and IaC, and supporting security/compliance and developer experience.
Top Skills:
Apollo GraphqlArgo CdAWSAws Secrets ManagerCircleCICloudFormationCloudfrontCloudwatchDatadogDockerEc2EcsEksFargateGithub ActionsGitopsGoGrafanaHelmIamIstioJavaJenkinsKubernetesKustomizeLambdaNewrelicOpsgeniePagerdutyPrometheusPythonRdsRoute53S3ServicenowSplunkTerraformTyk GatewayVaultVpc
Fintech • Financial Services
Leads SRE, DevSecOps, cloud engineering, observability, infrastructure automation, security automation, and CI/CD strategy across the Digital organization. Oversees system reliability, availability, scalability, performance, monitoring, incident management, risk mitigation, and operational efficiency. Builds and coaches engineering teams, establishes KPIs, partners with product and engineering leaders, and drives cloud-native delivery and continuous improvement using Azure, Terraform, Kubernetes, and related technologies.
Top Skills:
Agile ScrumApplication Performance Monitoring (Apm)AzureAzure NetworkingCi/CdCloud SecurityDevsecopsInfrastructure AutomationKubernetesObservabilitySite Reliability Engineering (Sre)TerraformTest Automation
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Financial Services
Migrate and maintain applications in Google Cloud, implement observability, monitor system health, support production on-call rotations, manage incidents, conduct post-incident reviews, and improve operational resiliency. The role drives automation, reduces toil, plans capacity, and ensures applications meet availability, performance, security, and maintainability standards.
Top Skills:
AutomationCloud InfrastructureGCPObservabilityProduction Systems
Reposted 15 Days AgoSaved
Artificial Intelligence • Automotive • Machine Learning • Software
Lead SRE ownership of ML platform SLOs and operational health for Ray on EKS and Databricks on EC2. Maintain observability with CloudWatch/Datadog, tune autoscaling and GPU scheduling, manage Databricks workspaces and IAM, optimize cost/capacity, codify infrastructure with Terraform and CI/CD, lead incident response and postmortems, perform security/OS maintenance, and participate in on-call rotation.
Top Skills:
Amazon LinuxAws Ec2Aws EksCi/CdCloudwatchDatabricksDatadogDockerDynamoDBEc2 SpotEcsGpu SchedulingIamKubernetesLambdaOn Demand Capacity ReservationsPythonRayRds/AuroraS3SqsTerraformUbuntuUnity Catalog
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills:
BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
Fintech • Analytics
Operate and improve an internal observability platform: monitor platform health, support incident response, define SLOs/SLIs/runbooks, automate GitOps workflows, enable teams with documentation and telemetry standards to improve reliability and reduce manual work.
Top Skills:
BigpandaCi/CdClickhouseCloudContainersCriblDatadogFlinkGitGitopsGrafanaInfrastructure-As-CodeLinuxNetworkingOpentelemetryRedis
Aerospace • Other
Build, operate, and scale mission-critical application infrastructure and tooling for vehicle and satellite software delivery. Manage infrastructure as code, improve observability, collaborate with engineers, participate in on-call rotation, perform incident response and postmortems, and provide end-user support to reduce build and test times.
Top Skills:
AnsibleBazelBuckC#C++ClickhouseDockerJavaScriptKubernetesKvmLinuxMakeMySQLPostgresPuppetPythonQemuTerraformVsphere
Software
Operate and improve enterprise IAM platforms for high availability and security. Automate infrastructure and identity workflows (Terraform, Ansible, Tines, scripting). Implement observability (metrics, logs, SLIs/SLOs) and lead incident response, runbooks, and change governance while collaborating across teams.
Top Skills:
AnsibleBashDatadogGrafanaIamPowershellPrometheusPythonSplunkTerraformTines
Information Technology
Lead cloud-native rearchitecture of a large monolithic system into microservices using AWS/Azure and Kubernetes. Design resilient, highly available infrastructure with Terraform, implement monitoring/alerting, solve complex cloud networking and security challenges, drive DevOps/CI-CD adoption, evaluate technologies, and mentor engineering teams.
Top Skills:
AksAWSAzureCloud NetworkingDevOpsDomain-Driven Design (Ddd)EksKubernetesLoad Balancing AlgorithmsTerraform
Big Data • Software
Maintain and improve service reliability and scalability by monitoring performance, troubleshooting production issues, implementing SLOs, performing capacity analysis, and developing automation and self-service tools. Participate in reliability testing, resilience evaluation, and release management while collaborating with teams to align monitoring and SLOs with user expectations.
Top Skills:
Aws AlbAws NlbAzure Application GatewayBgpDhcpDnsEigrpF5FirewallFlow LogsFortigateIds/IpsIp Addressing/SubnettingLoad BalancerMicro-SegmentationNaclsNatNetwork Monitoring PlatformsNginxOspfPacket AnalyzerPrivatelinkSecurity GroupsVnetVpcVpc PeeringVpnVpn TechnologiesWaf
Angel or VC Firm • Blockchain • Fintech • Cryptocurrency
Apply to join Galaxy Ventures' invite-only Talent Network for DevOps, SRE, QA, and Security professionals. Upon acceptance, your profile may be discreetly shared with portfolio companies for relevant roles, and you'll receive invitations to exclusive networking events. Participation is confidential and non-binding.
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Design, deploy, and maintain on-premises and cloud playout infrastructure for IP video distribution. Build automation, CI/CD pipelines, monitoring, and scalable fault-tolerant systems. Drive releases, troubleshoot broadcast incidents, mentor SREs, and provide 24/7 on-call support.
Top Skills:
AnsibleAWSAzureBashBroadcast TechnologiesCi/CdContainerizationGCPIp VideoJavaScriptKubernetesLinuxPerlPythonRubyStreamingTerraform
Artificial Intelligence • Cloud • Information Technology • Legal Tech • Productivity • Software
The Senior Site Reliability Engineer will focus on automating infrastructure, enhancing cloud resilience, supporting deployments, and mentoring teams in reliability best practices, while participating in on-call rotations.
Top Skills:
AzureBashCi/CdDockerGoGrafanaJavaKubernetesPowershellPrometheusPythonRubyTerraform
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Computer Vision • Hardware • Machine Learning • Robotics • Software
The role involves maintaining cloud infrastructure, collaborating with engineering teams, troubleshooting issues, deploying solutions, and ensuring system reliability.
Top Skills:
AnsibleC++GrafanaHelmKubernetesPagerdutyPythonTerraformTypescript
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results








_0.png)





.jpg)





.png)

.png)












