Top Site Reliability Engineer Jobs

Reposted 17 Days AgoSaved
Hybrid
O'Fallon, MO, USA
Senior level
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
The Senior Site Reliability Engineer will enhance service reliability, implement CI/CD using various tools, automate processes, and mentor junior resources.
Top Skills: ArtifactoryBitbucketCC++ChefGitGoJavaJenkinsMavenPerlPythonRuby
Reposted 17 Days AgoSaved
Remote or Hybrid
San Francisco, CA, USA
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
12 Days AgoSaved
In-Office
San Francisco, CA, USA
Entry level
Entry level
Artificial Intelligence • Big Data • Information Technology • Software • Analytics
Own the reliability, scalability, and operational health of an AWS and Kubernetes-based platform supporting connected sensors. Responsibilities include production troubleshooting, incident leadership, infrastructure management with Terraform, automation, CI/CD improvements, observability, service-level objectives, runbooks, and post-incident reviews. The role partners with application, AI, embedded systems, and fleet teams and participates in on-call support.
Top Skills: Amazon EksAWSBashCCi/CdDnsGitopsGoIamKubernetesLinuxNetworkingPythonRustTerraformVpn
12 Days AgoSaved
Remote or Hybrid
2 Locations
230K-250K Annually
Senior level
230K-250K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Software
Define reliability strategy and technical roadmaps; design scalable cloud infrastructure; improve observability, incident response, disaster recovery, automation, and deployment systems. Partner across engineering, security, data, AI, and product teams to strengthen resilience, compliance, and operational excellence. Lead architecture reviews, resolve complex production issues, establish SRE practices, and mentor engineers while reducing operational toil and improving platform scalability.
Top Skills: AutomationCi/CdCloud InfrastructureContainer OrchestrationContainersDisaster RecoveryDistributed SystemsInfrastructure As CodeMulti-Region ArchitectureNetworkingObservability
Reposted 12 Days AgoSaved
In-Office
Sunnyvale, CA, USA
145K-175K Annually
Senior level
145K-175K Annually
Senior level
Hardware • Semiconductor • Manufacturing
The Site Reliability Engineer will design, implement, and manage reliable infrastructure and services, ensuring operational excellence and uptime.
Top Skills: AWSBashDockerGrafanaKubernetesLinuxAzureOpenshiftPrometheusProxmoxPythonVmware Vsphere
Reposted 12 Days AgoSaved
Remote
United States
120K-165K Annually
Senior level
120K-165K Annually
Senior level
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills: AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Reposted 12 Days AgoSaved
In-Office
Reston, VA, USA
137K-244K Annually
Senior level
137K-244K Annually
Senior level
Cloud • Fintech • HR Tech
Work as an SRE on the Analytics Delivery Engineering team to design, automate, and maintain a secure, highly available Kubernetes platform. Build CI/CD pipelines, infrastructure-as-code, observability, and automation for Prism Analytics in GovCloud. Troubleshoot production incidents, participate in on-call rotations, collaborate across teams, and ensure compliance and security for federal deployments.
Top Skills: SparkArgo CdAWSAws GovcloudCi/CdDockerGoGrafanaKubernetesObservabilityPrism AnalyticsPrometheusPythonTerraformTracing
13 Days AgoSaved
In-Office
New York, NY, USA
150K-160K Annually
Senior level
150K-160K Annually
Senior level
Natural Language Processing • Software • Conversational AI
Design, deploy, and manage resilient GCP network architectures, including VPCs, Interconnects, load balancing, and DNS. Lead Terraform-based infrastructure automation, implement network security policies, optimize cloud networking costs, and resolve connectivity and performance issues. Drive SRE practices such as capacity planning, error budgets, monitoring, and alerting. Partner with SRE, platform, and engineering teams to onboard services and maintain high availability, reliability, and security across distributed systems.
Top Skills: BashCloud InterconnectDnsGoGoogle Cloud Platform (Gcp)Infrastructure As Code (Iac)LinuxLoad BalancingPythonTerraformVpc
13 Days AgoSaved
In-Office or Remote
2 Locations
142K-268K Annually
Expert/Leader
142K-268K Annually
Expert/Leader
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills: Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
13 Days AgoSaved
In-Office or Remote
Ann Arbor, MI, USA
Expert/Leader
Expert/Leader
Big Data
Technical leader responsible for reliability, automation, scalability, and infrastructure across a cloud platform. Designs and operates Kubernetes-based systems, Infrastructure as Code, CI/CD, observability, incident response, internal platforms, and agentic AI/LLM workloads. Leads reliability practices, complex troubleshooting, operational automation, on-call improvements, architecture influence, cross-team initiatives, and mentorship while remaining hands-on with code, infrastructure, and incidents.
Top Skills: Agent Orchestration FrameworksAWSAzureCi/CdDockerElasticsearchFluxcdGCPGithub ActionsGitopsGoGrafanaHelmInfrastructure As CodeJavaJenkinsKafkaKubernetesLinuxLlm GatewaysLokiObservabilityOpentofuPostgresPrometheusPythonSentrySignozTcp/IpTerraform
Reposted 13 Days AgoSaved
In-Office
30339, Atlanta, GA, USA
Senior level
Senior level
Retail
Lead Dynatrace and Azure observability efforts: design and implement telemetry, DQL/KQL analytics, dashboards, alerts, and automation; enable mobile and .NET observability; troubleshoot Azure Functions, APIs, and distributed systems; perform RCA, reduce incident recurrence, and drive SRE best practices and automation.
Top Skills: .NetAndroidAnsibleApplication InsightsAzure Api Management (Apim)Azure App ServicesAzure FunctionsAzure Kusto Query Language (Kql)Azure Log AnalyticsAzure MonitorCC++ChefCi/CdDatabricksDevOpsDynatraceDynatrace Davis AiDynatrace Query Language (Dql)GoiOSJavaJenkinsOpentelemetryPuppetPythonSccmServicenowSQLTerraform
13 Days AgoSaved
In-Office
Headquarters, AZ, USA
Internship
Internship
Automotive
Support 24/7 site reliability operations by monitoring global data centers and applications, building Python and Ansible automation, troubleshooting infrastructure issues, managing incident response, performing security patching and failover testing, and contributing to root cause analysis and operational improvements. The intern will collaborate with infrastructure, DevOps, networking, systems, database, and product teams while maintaining documentation, reports, procedures, and training materials.
Top Skills: AnsibleAWSDatadogGCPKubernetesLinuxPythonVmware VsphereWindows
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 13 Days AgoSaved
In-Office
Albany, NY, USA
100K-150K Annually
Senior level
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
Build and maintain enterprise digital experiences on AEM (Sites, Assets, Cloud). Hands-on development of components, templates, workflows, integrations, front-end experiences, GraphQL/headless delivery, CI/CD, performance tuning, and stakeholder communication.
Top Skills: Adobe Experience Manager (Aem)Adobe Experience PlatformAem As A Cloud ServiceAem AssetsAem SitesCi/CdContent FragmentsDispatcherEdge Delivery ServicesFranklinGraphQLHeadless DeliveryHlxHtlJavaJavaScriptMagentoModern Js Build ToolsOsgiSalesforce Commerce CloudSlingSpa Editor
13 Days AgoSaved
In-Office
Alpharetta, GA, USA
87K-144K Annually
Senior level
87K-144K Annually
Senior level
Information Technology • Legal Tech • Analytics
Lead security-focused site reliability engineering across cloud and on-premises platforms. Build observability, SIEM integrations, compliance-as-code, policy enforcement, infrastructure hardening, secrets management, vulnerability scanning, and AI-enhanced security operations. Support Azure, AWS, Kubernetes, CI/CD, incident response, disaster recovery, and 24/7 production systems while aligning controls with NIST, ISO 27001, and SOC 2. Collaborate across engineering and governance teams and train staff on security best practices.
Top Skills: ArgocdAWSBashChainguardCheckovCis BenchmarksDlpGithub ActionsGrafanaIso 27001KubernetesLokiMfaAzureNistOidcOtelPowershellPrometheusPythonQualysSIEMSnykSoc 2TerraformTrufflehogVaultWizZero Trust
13 Days AgoSaved
In-Office
San Jose, CA, USA
105K-155K Annually
Junior
105K-155K Annually
Junior
Software
Build and maintain SRE microservices supporting a global GPU infrastructure platform. Implement GitOps, declarative configuration, and CI/CD automation; monitor metrics, logs, and traces; support operational readiness through on-call participation, runbooks, and incident reviews; and write unit, integration, and end-to-end tests. Work with senior engineers to deliver production-ready platform features while maintaining service reliability and infrastructure consistency.
Top Skills: AnsibleArgocdBashCi/CdCudaDockerFluxGitopsGoGrafanaHelmJavaKubernetesLinuxLokiNvidia DcgmOpentelemetryPrometheusPythonRustTerraform
13 Days AgoSaved
In-Office
Salt Lake City, UT, USA
Senior level
Senior level
Cloud • Software • Analytics
Lead the vision, strategy, and roadmap for core enterprise SaaS platform components. Gather technical requirements, prioritize features, maintain technical documentation, and oversee engineering execution. Monitor product performance and market trends while identifying enhancement opportunities. Partner with marketing and sales on technical positioning and go-to-market strategies. The role requires strong product management, engineering, stakeholder collaboration, Agile, Lean, and regulated-industry experience.
Top Skills: AgileAPIsData AnalyticsLeanSaaSSoftware Validation
19 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Build and operate scalable cloud-native microservices and Kubernetes infrastructure, improve CI/CD and developer workflows, engineer and run an autonomous coding-agent orchestration platform, integrate and operationalize AI services with guardrails, implement observability and incident response, and collaborate with product and engineering teams to ensure secure, reliable production systems and cost optimization.
Top Skills: Ai ApisAutonomous AgentsAWSCi/CdGCPGoJavaJavaScriptKubernetesMonitoringObservabilityPythonRestful ApisRustSdksSQLTypescriptWorkflow Orchestration
Reposted 13 Days AgoSaved
In-Office
Birmingham, AL, USA
Senior level
Senior level
Automotive • Hardware • Logistics
The Site Reliability Engineer III enhances system reliability by building automation and supporting large-scale systems, ensuring critical platforms function optimally.
Top Skills: APIsAzure DevopsDynatraceGoogle Cloud PlatformGrafanaHTTPJavaKubernetesMicroservicesPrometheusTerraform
Reposted 13 Days AgoSaved
In-Office
Boston, MA, USA
160K-225K Annually
Senior level
160K-225K Annually
Senior level
Hardware • Quantum Computing
Lead integration, maintenance, and automation of heterogeneous hardware and software control systems for quantum computers. Manage networked lab infrastructure, CI/CD pipelines, observability, and provisioning. Support incident response, testing, and orchestration, collaborating with software, hardware, and test teams to ensure reliability and operational readiness of development and production environments.
Top Skills: AnsibleBashCi/CdDebianDhcpDnsDockerElkGitGitlab CiGoGrafanaHardware-In-The-Loop (Hil)JenkinsKubernetesLanLogging SystemsPrometheusPythonRack-Mount ServersRed HatRoutersSwitchesTcp/IpTerraformUbuntuVlanWanWindows
Reposted 13 Days AgoSaved
In-Office
Austin, TX, USA
167K-204K Annually
Senior level
167K-204K Annually
Senior level
Automotive • Information Technology • Logistics • Software
Lead Site Reliability Engineer implements IaC and automation, builds observability (SLIs/SLOs, dashboards, alerting), manages incident response, runbooks, gamedays, postmortems, and drives SRE/DevOps best practices, AppSec integration, testing, and CI/CD improvements across teams.
Top Skills: AppsecAWSAws CloudformationC#Ci/CdCloudsploitCloudwatchData TheoremDatadogGrafanaIacInfrastructure As CodeJavaNewrelicPythonTerraformVeracode
Reposted 13 Days AgoSaved
Remote
USA
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills: AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Reposted 13 Days AgoSaved
In-Office or Remote
The Center, IN, USA
Expert/Leader
Expert/Leader
Edtech • Information Technology • Software
Lead infrastructure, reliability, and observability across multi-cloud environments. Improve CI/CD, IaC standards, staging parity, Kubernetes operations, monitoring and SLOs, incident response, and platform modernization while partnering with engineering teams.
Top Skills: Ai-Assisted Development Tools (Claude CodeAutoscalingAWSCi/Cd PipelinesCodex)Event-Driven ArchitecturesGCPIncident ManagementInfrastructure-As-CodeKubernetesMonitoringObservabilityPythonQueue-Based ArchitecturesRuby On RailsSlo FrameworksTerraform
Reposted 13 Days AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Software • Generative AI
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating cloud-native systems, owning production reliability and incident response, building observability and automation, defining SLOs/SLIs, driving postmortems, and partnering with product and engineering teams to improve operational maturity.
Top Skills: AWSAzureGCPGoJavaKubernetesPython
Reposted 13 Days AgoSaved
In-Office
Sunnyvale, CA, USA
170K-200K Annually
Senior level
170K-200K Annually
Senior level
Security • Software • Cybersecurity
Hands-on Site Reliability Engineer responsible for building and maintaining cloud infrastructure, CI/CD pipelines, observability (logging/monitoring/tracing), automation, and security best practices. Manage datacenter resources, troubleshoot clusters and services, collaborate with engineering teams for deployments, and participate in on-call incident response to ensure high availability and performance.
Top Skills: AnsibleArgocdBashChefDatadogElkGitlab CiGoGrafanaJenkinsKubernetesLinuxPrometheusPythonRancher
Reposted 13 Days AgoSaved
In-Office or Remote
2 Locations
76K-136K Annually
Mid level
76K-136K Annually
Mid level
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills: Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account