Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Fintech • Analytics
Maintain SLOs and improve availability, latency, and system health for cloud-hosted services. Design automation and IaC for AWS/Azure, support cloud migration, configure observability (Datadog, CloudWatch, Azure Monitor), participate in on-call rotations and incident response, and partner with development teams to improve reliability, observability, and CI/CD pipelines.
Top Skills:
AksAws CloudformationAws CloudwatchAws CodepipelineAws Secrets ManagerAws Well-Architected FrameworkAzure (Services)Azure Architecture CenterAzure Arm TemplatesAzure DevopsAzure MonitorAzure SqlDatadogDockerDynamoDBEc2EksGitIamKmsKubernetesLambdaPythonRdsS3ShellTerraformVpc
Other
Design, build, and maintain scalable, secure cloud infrastructure and automation. Implement IaC and CI/CD, optimize customer-facing platforms, manage CDNs/DNS/SSL, improve observability and incident response, mentor engineers, and drive infrastructure strategy and reliability improvements.
Top Skills:
Akamai CdnAnsibleArgo CdAWSAzureAzure DevopsCdnChefCi/CdCloudFormationComptia Security+DnsGitGCPHTTPHttpsIncident ManagementLinuxMonitoringObservabilityPuppetSsl/TlsTerraformWindows
Cloud • Information Technology • Internet of Things • Professional Services • Software
Ensure production database services are scalable, resilient, high-performing, and secure. Participate in on-call rotation, monitoring, alerting, incident response, disaster recovery drills, and post-mortems. Automate operational tasks, design reliability/security controls, build observability tooling, capacity forecasts, and collaborate with developers and security to improve hybrid cloud vector and NoSQL database platforms.
Top Skills:
AWSAzureCi/CdGCPGitGitJenkinsKubernetesLangchainMilvusMongoDBMySQLNoSQLOpenai ApiOpenshiftPineconePostgresPythonVector DatabasesWeaviate
Information Technology
Lead SRE responsible for reliability, performance, scalability, observability, automation, and incident response across cloud and air-gapped environments. Drive RCA, capacity planning, reliability standards, and tooling while partnering with DevOps, infrastructure, and security teams to reduce operational risk and support highly available services.
Top Skills:
AWSElk StackGrafanaKubernetesLinuxPrometheusPythonTerraformTerragrunt
Cybersecurity
Own reliability, observability, and delivery for a cloud-native, multi-tenant Kubernetes platform. Build CI/CD, GitOps, IaC, and AI-first automation for incident response and LLM infrastructure. Instrument SLOs and tracing, harden secrets and supply-chain security, troubleshoot production issues, and mentor engineers in AI-driven operations.
Top Skills:
Agentic OrchestrationCdcCi/CdClaude CodeCloud (Major Providers)CursorData PipelinesDistributed TracingDockerGitopsGoInference GatewaysInfrastructure-As-CodeKubernetesLlmsObservabilityPost-Quantum CryptographyPythonSecrets ManagementService MeshSlosStreamingSupply-Chain SecurityWindsurf
AdTech • Big Data • Cloud • Marketing Tech • Software • Analytics
Lead SRE efforts to improve security, reliability, cost efficiency, and observability. Build automation, CI/CD, agentic AI platforms (MCPs), and tooling for capacity planning, incident response, and self-service. Evangelize SecDevOps and zero-trust designs across product and platform teams.
Top Skills:
Ai Agentic FrameworksArgocdAWSBashCi/CdDevsecopsDockerEksGoGrafanaKubernetesLinuxLokiMcpNew RelicPrometheusPythonSamTerraformZero Trust
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead Site Reliability Engineering for Mastercard Business Operations: ensure production readiness, reliability, scalability, and performance of applications. Implement observability, automation, CI/CD, containerization, and cloud infrastructure best practices. Support incident response, capacity planning, troubleshooting, and risk/compliance activities while mentoring developers and promoting developer-run ownership and operational standards.
Top Skills:
AWSAzureBashCi/CdContainerizationGCPGoLinuxNetworkingObservabilityOrchestrationPythonUnix
Financial Services
Deploy, manage, and optimize the Splunk platform and observability tooling. Onboard data sources, build SPL queries and dashboards, monitor platform health, tune performance, respond to incidents/on-call rotations, support Java and .NET applications, perform root cause analysis, and collaborate across engineering, security, and operations to maintain reliable, secure enterprise services.
Top Skills:
.NetAnsibleApache Http ServerAppdynamicsApplication Performance ManagementArtifactoryDockerGitGitflowIbm AixIbm Secure Directory ServerIbm Security Access Manager (Isam)JavaJava Enterprise EditionJenkinsJvmKubernetesMicrosoft Windows Server 2008Microsoft Windows Server 2012OpenshiftPythonQuayRed Hat LinuxSearch Processing Language (Spl)SnmpSplunkSplunk Common Information Model (Cim)
Information Technology • Cybersecurity • Defense • Automation
Design, build, and maintain secure, highly available Azure cloud-native platforms and CI/CD pipelines for classified mission systems. Implement Infrastructure as Code, automation, monitoring, and security controls while supporting hybrid Windows environments and platform reliability.
Top Skills:
Arm TemplatesAzureAzure ComputeAzure DevopsAzure IdentityAzure Kubernetes ServiceAzure MonitorAzure NetworkingAzure StorageBicepC#Ci/CdDevsecopsDockerGitInfrastructure As CodeKubernetesLog AnalyticsPowershellPythonTerraformWindows Server
Reposted 18 Days AgoSaved
Aerospace • Defense • Manufacturing
Lead and build the deployment engineering function to operate mission-critical software in accredited, air-gapped, and high-side environments. Own full deployment lifecycle across cloud, on-prem, and disconnected networks; manage Kubernetes/OpenShift and Linux infrastructure; build CI/CD and IaC workflows; integrate security tooling; diagnose and prevent production issues; produce ATO-related artifacts and maintain compliance.
Top Skills:
AlertmanagerAWSAzureBashDockerGCPGitlab CiGoGrafanaGroovyHelmJavaJenkinsKubernetesOpenshiftPagerdutyPodmanPrometheusPythonRhelRubyService MeshSplunkTerraform
Hardware • Other • Energy
Maintain and monitor production systems for availability and performance; lead incident response and postmortems; implement observability, alerting, and automated remediation; optimize distributed systems (AKKA.NET) and PostgreSQL; build CI/CD pipelines and infrastructure-as-code.
Top Skills:
Akka.NetAWSAzureAzure DevopsAzure PipelinesBashC#DatadogDockerElkGCPGitGithub ActionsGitlabGitlab CiGrafanaKubernetesOpentelemetryPhobosPostgresPowershellPrometheusPythonTerraform
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills:
AWSKubernetesPostgresTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Reposted 18 Days AgoSaved
Aerospace • Information Technology • Professional Services • Security • Software
Maintain and improve reliability, scalability, and performance of enterprise infrastructure across global sites. Implement automation and infrastructure-as-code, build monitoring and observability, perform RCA and incident response, support patching and RMF changes, integrate new capabilities, and maintain operational documentation and ITIL/ITSM processes to ensure mission-ready, high-availability environments.
Top Skills:
AnsibleElkNagiosPowershellPythonScomSolarwindsSplunkTerraform
Big Data • Energy • Big Data Analytics
The Staff Site Reliability Engineer will lead in designing and maintaining cloud infrastructure on GCP, drive IaC strategy, manage Kubernetes operations, ensure security compliance, and mentor engineers.
Top Skills:
BashGoGoogle Cloud PlatformGrafanaKubernetesOpentelemetryPostgresPythonTerraform
Hardware • Manufacturing
Operate and harden a multi-cloud microservices platform: deploy on Kubernetes, run load/chaos tests, build observability, automate with scripts, define SLO/SLA, ensure security/compliance, participate in incident response, disaster recovery, on-call rotation, and mentor junior team members.
Top Skills:
AWSAzureBashGCPGoHpaJavaJvmKubernetesMicroservicesOciPowershellPython
Reposted 18 Days AgoSaved
Fintech • Financial Services
Lead SRE technical strategy and architecture for highly available, scalable enterprise platforms. Build automation, observability, and incident response practices; mentor senior engineers; drive capacity planning, production reliability, and adoption of SRE best practices across cloud and on-prem environments.
Top Skills:
AnsibleAWSBigQueryChefCloudFormationDatadogDockerElasticsearchElk StackGCPGitlabGoGrafanaJavaJenkinsKafkaKubernetesLinuxMavenPagerdutyPrometheusPrompt EngineeringPuppetPythonRetrieval-Augmented Generation (Rag)Terraform
Cloud • Security
Build and operate the production platform (Kubernetes, AWS, IaC, CI/CD, observability), automate self-service deployment, embed security and secrets management, run and modernize on-call, drive cost efficiency, mentor teammates, and maintain runbooks and post-incident reviews.
Top Skills:
AWSBashCi/CdClaudeGitGrafanaKubernetesLinuxPrometheusPythonSaltTerraform
Aerospace • Defense
Own and operate Loft's cloud and hybrid network infrastructure: design architectures, secure connectivity, manage network-as-code (IaC), define SLOs and observability, automate toil, and participate in SRE incident response and platform reliability.
Top Skills:
ArgocdCi/CdDnsDockerFirewallingFluxcdGCPGitGitopsGrafanaHybrid ConnectivityK8SKubernetesPeeringSdnSite-To-Site RoutingSoftware-Defined NetworkingTerraformVpcVpn
Gaming
The role involves ensuring production quality, owning system reliability, and participating in decision-making. Responsibilities include incident response and lifecycle management in cloud gaming technologies.
Top Skills:
BashC++ElasticsearchGoIstioJavaKafkaKong Api GatewayKubernetesKumaLinkerdMongoDBMySQLPostgresPythonRedisRust
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills:
ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
Information Technology • Professional Services
Provide senior SRE expertise to improve reliability, scalability, performance, and resilience of a cloud-hosted geospatial platform. Design monitoring/observability, automate deployments, support incident response, optimize capacity and performance, and collaborate across DevSecOps, Kubernetes, database, and support teams in a mission-focused DoD environment.
Top Skills:
Alerting ToolsArcgis EnterpriseAw S Cloud OneAWSCi/CdContainerized SystemsEsriInfrastructure-As-CodeKubernetesLinuxLogging ToolsMonitoring ToolsRmfScriptingStig
Information Technology • Insurance • Software
Own and operate production services end-to-end to ensure reliability, scalability, performance, and operational health. Define SLIs/SLOs, perform incident response and root cause analysis, build automation and self-healing, manage production changes, and collaborate with engineering, product, and operations teams to improve system design and observability.
Top Skills:
.NetAWSC#Ci/CdInfrastructure As CodeJavaKubernetesLinuxPythonReactRelational DatabasesWindows
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills:
Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills:
AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
Cloud
Design, build, and maintain secure, air-gapped cloud platform services and CI/CD pipelines for Okta Federal. Operate mission-critical infrastructure, monitor SLOs/SLIs, run incident response and POA&M remediation, and support Authority to Operate activities. Advocate SRE/DevOps practices across teams while working autonomously in secure facilities.
Top Skills:
Air-Gapped EnvironmentsAmazon CloudwatchAws Transit GatewayAws VpcBgpCi/CdEcs FargateEksGrafanaIpsecPythonSplunkTerraformVpc Endpoints
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results
.jpg)


































