Top Site Reliability Engineer Jobs

23 Days AgoSaved
In-Office or Remote
3 Locations
405K-485K Annually
Senior level
405K-485K Annually
Senior level
Artificial Intelligence • Natural Language Processing • Generative AI
Lead deployment and runtime operations for ML safety systems: configure and verify safeguards across platforms, run canary rollouts and validations, automate launch runbooks into pipelines, maintain a provenance-backed safeguards registry, and participate in on-call and incident response for safety-critical model releases.
Top Skills: AWSAws BedrockCanary DeploymentsCi/Cd PipelinesClaudeConfig ManagementGCPGcp VertexLlm InferencePythonRustTransformer-Based Models
29 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills: Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Reposted 23 Days AgoSaved
Remote
United States
206K-263K Annually
Expert/Leader
206K-263K Annually
Expert/Leader
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills: Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Reposted 23 Days AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills: AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
YesterdaySaved
Hybrid
Bellevue, WA, USA
160K-210K Annually
Senior level
160K-210K Annually
Senior level
Artificial Intelligence • Marketing Tech
Own and scale Cognitiv’s AWS infrastructure while evaluating architecture, networking, security, scalability, and service management. Lead improvements in deployments, monitoring, disaster recovery, and infrastructure-as-code practices. Support co-located Equinix datacenter deployments and hybrid cloud operations alongside a datacenter-focused SRE. Collaborate with engineering and product teams, provide multi-datacenter coverage, and help establish long-term service management best practices.
Top Skills: AnsibleAWSBashDatadogEc2EquinixKubernetesPrometheusPythonTerraform
24 Days AgoSaved
In-Office
Taylor, TX, USA
90K-115K Annually
Mid level
90K-115K Annually
Mid level
Computer Vision • Hardware • Mobile • Software • Semiconductor
Ensure 24x7 reliability and scalability of AMHS control software in a semiconductor fab. Build and maintain back-end data pipelines, observability (Prometheus/Grafana), and automation tooling. Troubleshoot software, IPC/RPC network issues, and optimize middleware (RabbitMQ, Redis). Modernize infrastructure using Python, Java, Docker, and Kubernetes to reduce toil and prevent production downtime.
Top Skills: DockerETLGrafanaIpcJavaKubernetesLinux/UnixMesNosql (Nosql Databases)PrometheusPythonRabbitMQRedisRpcSecs/GemSql (Relational Databases)
Reposted YesterdaySaved
Remote or Hybrid
US
138K-221K Annually
Senior level
138K-221K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills: .NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
24 Days AgoSaved
Remote
United States
145K-200K Annually
Senior level
145K-200K Annually
Senior level
Software
Lead SRE to define strategy and roadmap for reliability, scalability, observability, and automation across cloud and hybrid environments. Design and operate containerized production workloads, infrastructure-as-code, monitoring/alerting, incident management, and compliance for regulated domains. Mentor SREs, partner with security and product teams, manage cloud costs and capacity, and build a developer platform to improve delivery and production stability.
Top Skills: AWSAws MarketplaceAzureAzure MarketplaceGCPGoogle Cloud MarketplaceGrafanaKubernetesPrometheusTerraform
24 Days AgoSaved
In-Office
Great Neck, NY, USA
90-115 Annually
Senior level
90-115 Annually
Senior level
Healthtech • Software
Own reliability, security, and continuity of Medgen production environments. Manage Windows and Linux servers, HA configurations, backups, DR, deployments, and CI/CD automation. Coordinate vulnerability assessments, incident response, cloud/hybrid evaluations, and infrastructure modernization.
Top Skills: .NetAWSAzureBackupsCi/CdDisaster RecoveryDnsFirewallsGCPIisJavaScriptLinuxMicrosoft TfsNetwork SegmentationSQLTcp/IpVpnsWindows Server
24 Days AgoSaved
In-Office
Great Neck, NY, USA
90-115 Annually
Senior level
90-115 Annually
Senior level
Healthtech • Software
Own reliability, security, and continuity of Labgen production: manage Windows/Linux servers, HA, networking, backups/DR, production deployments, CI/CD, monitoring, cloud evaluation, incident response, vendor relations, and compliance.
Top Skills: .NetApacheAWSAzureBashCentosCertificates/PkiCvsDatadogDnsElkFirewallsGCPGitGrafanaJavaScriptJenkinsLinux (RhelPrometheusPythonSentineloneSql/Relational DatabasesTcp/IpUbuntu)VpnWindows Server
YesterdaySaved
In-Office
Los Angeles, CA, USA
108K-162K Annually
Senior level
108K-162K Annually
Senior level
Fintech • Financial Services
Operates and improves high-traffic, business-critical cloud and network systems. Designs scalable network solutions, manages performance and troubleshooting, automates deployment and monitoring, coordinates router and switch installations, and tests redundancy, resilience, and failover. Partners with development teams to improve operability, supports project planning and vendor comparisons, provides training, and participates in on-call coverage.
Top Skills: Cloud ComputingIpRoutersSwitchesVoip
YesterdaySaved
Remote
United States
Senior level
Senior level
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills: Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
YesterdaySaved
In-Office
3 Locations
153K-192K Annually
Senior level
153K-192K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
Designs and matures site reliability engineering capabilities across cloud and on-premises platforms. Defines SLIs, SLOs, error budgets, observability standards, and production readiness requirements. Automates toil reduction, remediation, infrastructure provisioning, CI/CD, and operational workflows. Leads incident triage, root-cause investigations, capacity planning, resiliency initiatives, AIOps adoption, and platform health monitoring. Partners with development, infrastructure, security, and governance teams to improve reliability, scalability, compliance, and operational excellence.
Top Skills: AiopsAnsibleCi/CdCloud PlatformsContainerizationEvent-Driven RemediationGenaiIamInfrastructure As Code (Iac)Infrastructure AutomationLogsMetricsObservabilityOrchestration PlatformsPolicy-As-CodeSynthetic MonitoringTelemetryTerraformTraces
YesterdaySaved
In-Office
Chicago, IL, USA
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Financial Services
Define and implement the company’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerting, reliability patterns, and observability standards. Embed circuit breakers, retries, bulkheads, and load shedding into Ruby, Java, and Elixir services. Extend Prometheus, Honeycomb, and OpenTelemetry monitoring, support scaling on HashiCorp Nomad, mentor engineers, and establish blameless incident review practices.
Top Skills: ConsulElixirFlow AnalysisGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Reposted 24 Days AgoSaved
Remote
United States
147K-168K Annually
Senior level
147K-168K Annually
Senior level
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills: AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Reposted 24 Days AgoSaved
Remote
United States
235K-275K Annually
Expert/Leader
235K-275K Annually
Expert/Leader
Legal Tech • Software
Senior technical leader for SRE driving observability, platform infrastructure, SLIs/SLOs, incident response, automation, and self-service platform capabilities. Shapes reliability strategy, mentors engineers, and ensures production-scale operational excellence.
Top Skills: AiopsBashDatadogGoInfrastructure As CodeKubernetesNew RelicObservabilityPython
24 Days AgoSaved
Remote
30 Locations
Mid level
Mid level
Information Technology
Maintain and optimize mission-critical bare-metal infrastructure, build automation for web2/web3 deployments, perform R&D and monitoring for blockchain validator/RPC/operator nodes, ensure SLIs/SLOs, and participate in on-call rotation.
Top Skills: AnsibleBare-MetalDockerGCPGithub ActionsGoHaproxyKubernetesOracle CloudPostgresPythonTerraform
24 Days AgoSaved
In-Office
Austin, TX, USA
Senior level
Senior level
Artificial Intelligence • Big Data • Information Technology • Security • Software
Design, build, and maintain cloud infrastructure and CI/CD for a high-availability telecommunications product. Define SLOs/SLIs, manage incident response and on-call rotations, implement observability, perform performance and capacity planning, run blameless postmortems, and collaborate with security teams to ensure compliance and access control.
Top Skills: AnsibleAWSDatadogDockerGCPGitlabHelmJavaJenkinsKubernetesNoSQLTerraform
Reposted 24 Days AgoSaved
Hybrid
Camden, NJ, USA
130K-150K Annually
Mid level
130K-150K Annually
Mid level
Information Technology • Logistics • Transportation • Analytics • Business Intelligence • 3PL: Third Party Logistics • Industrial
As a Site Reliability Engineer, you'll enhance reliability for Phenix WMS and automation systems, focusing on incident reduction and system health through observability and automation. Responsibilities include defining SLIs and SLOs, participating in incident response, and testing disaster recovery plans.
Top Skills: AnsibleAzureBashCi/CdKubernetesPowershellPythonTerraform
Reposted 24 Days AgoSaved
In-Office
Hawthorne, CA, USA
125K-175K Annually
Mid level
125K-175K Annually
Mid level
Aerospace • Other
The Site Reliability Engineer, GNC at SpaceX oversees mission-critical GNC products, operates servers, maintains HPC clusters, and enhances services and infrastructure to support space operations.
Top Skills: AnsibleBazelDockerGradleKubernetesLinuxMakeNpmPipPuppetPythonTerraformVagrant
Reposted 24 Days AgoSaved
In-Office
Hawthorne, CA, USA
125K-175K Annually
Junior
125K-175K Annually
Junior
Aerospace • Other
Manage and design server, HPC, storage, and networking infrastructure (including InfiniBand) to support propulsion engineering workflows. Integrate and optimize engineering applications (ANSYS, StarCCM+), automate deployments, troubleshoot performance bottlenecks, and coordinate with IT, facilities, and engineering teams to scale compute resources for rocket engine development.
Top Skills: AnsibleAnsysBashDockerEnterprise NetworkingHpcInfinibandKubernetesLinuxPuppetPythonStarccm+VirtualizationWindows Server
Reposted 24 Days AgoSaved
In-Office
San Francisco, CA, USA
175K-250K Annually
Mid level
175K-250K Annually
Mid level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
The Site Reliability Engineer will ensure the reliability and performance of AI infrastructure, build core systems, handle incident response, and develop automation tools.
Top Skills: AWSDatadogElkGCPGithub ActionsGitlab CiGoGrafanaJenkinsKubernetesLinuxPrometheusPulumiPythonRustTerraform
Reposted 24 Days AgoSaved
In-Office or Remote
7 Locations
Senior level
Senior level
Database
Embed with service teams to define SLIs/SLOs and error budgets, run Operational Readiness Reviews, improve incident-to-improvement pipelines, advise on resilience and architecture, reduce operational toil through automation, and shape org-wide on-call practices and operational maturity.
Top Skills: AWSCdkGrafanaKubernetesOpentelemetryPostgresPulumiTerraformVictoriametrics
Reposted 24 Days AgoSaved
In-Office
Irving, TX, USA
96K-145K Annually
Senior level
96K-145K Annually
Senior level
Fintech • Financial Services
Provide production support and SRE for business-critical digital asset platforms (wallets, blockchain-connected systems, transaction processing). Lead incident management, troubleshoot high-availability distributed systems, perform system-level Linux debugging, use monitoring/observability tools, automate tasks via scripting, investigate data with SQL, and communicate with technical and business stakeholders.
Top Skills: BlockchainDistributed SystemsElkGrafanaJavaKubernetesLinuxNode OperationsOpenshiftPrometheusPythonSplunkSQLTransaction ProcessingWallets
Reposted 24 Days AgoSaved
Hybrid
2 Locations
Mid level
Mid level
Fintech • Financial Services
The Site Reliability Engineer will support cloud infrastructure, automate deployments, and ensure operational efficiency and governance across public cloud platforms.
Top Skills: AnsibleAWSAzureAzure CliAzure FunctionsAzure Kubernetes ServiceCosmodbGCPGitJenkinsKubernetesLinuxPowershellTerraformWindows
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account