Top Site Reliability Engineer Jobs

Reposted 24 Days AgoSaved
In-Office
Los Angeles, CA, USA
140K-199K Annually
Senior level
140K-199K Annually
Senior level
Fintech • Financial Services
Lead Site Reliability Engineer responsible for ensuring system reliability, scalability, and performance. Develop automated deployment strategies, maintain monitoring/observability, define SLIs/SLOs, collaborate with cross-functional teams, drive reliability best practices, participate in on-call incident response, and improve delivery through automation and training.
Top Skills: AWSAzureBashGCPPowershellPython
Reposted 24 Days AgoSaved
In-Office
New York, NY, USA
147K-310K Annually
Expert/Leader
147K-310K Annually
Expert/Leader
Fintech • Financial Services
The Director of Splunk Platform Engineering & SRE owns the enterprise Splunk platform, drives incident resolution, optimizes systems, and mentors engineers, focusing on automation and performance.
Top Skills: AnsibleGitGoJavaKubernetesLinux/UnixMoogPrometheusPythonSplunk
Reposted 24 Days AgoSaved
In-Office
Westminster, CO, USA
106K-145K Annually
Mid level
106K-145K Annually
Mid level
Hardware • Information Technology • Other • Software • Analytics
Architect and operate ML/agent pipelines and infrastructure, deploy and monitor models at scale, pioneer MLOps/Agent Ops best practices, collaborate with domain experts, and test/optimize ML systems for production reliability and cost efficiency.
Top Skills: Bash ScriptingContainerization (E.G.Docker)Git/GithubLinuxModel VersioningMonitoringNumpyPandasPythonPyTorchScikit-Learn
2 Days AgoSaved
In-Office or Remote
Location, WV, USA
Senior level
Senior level
Healthtech • Telehealth
Lead reliability engineering for mission-critical Azure healthcare workloads. Define and manage SLIs, SLOs, and error budgets; improve observability, scalability, resilience, disaster recovery, and cloud operations. Configure monitoring and alerting, conduct incident response and postmortems, automate remediation, support security and compliance controls, and partner with engineering, product, security, and operations teams. Mentor SRE staff and help establish reliability practices across the organization.
Top Skills: AnsibleApplication InsightsAWSAzure Chaos StudioAzure Kubernetes Service (Aks)Azure MonitorBicepChaos MeshDatadogDynatraceElastic/ElkGoGrafanaGremlinLog AnalyticsLogicmonitorAzureNew RelicOpentelemetryPowershellPrometheusPythonTerraform
2 Days AgoSaved
In-Office
Santa Clara, CA, USA
168K-334K Annually
Senior level
168K-334K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, deploy, and optimize on-premises HPC storage infrastructure integrated with cloud environments. Build automation, monitoring, alerting, and self-service tools for large-scale distributed storage systems. Manage enterprise NAS, S3, and parallel filesystems; evaluate technologies; resolve performance bottlenecks; document procedures; and collaborate with engineering teams on infrastructure requirements, application deployment, and resource utilization.
Top Skills: AWSAzureBashCloudianDistributed StorageDockerElasticsearchEnterprise NasGCPGoGpfsGrafanaHpcInfinibandKibanaKubernetesLsfLustreMesosphere DcosMinioNetappPbsPrometheusPure StoragePythonRdmaRoceS3SlurmSplunkZabbix
Reposted 2 Days AgoSaved
In-Office
2 Locations
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve Microsoft Defender services in US Government clouds: lead on-call incident response, automate deployments and monitoring, build scalable reliability tooling, ensure security/compliance, run postmortems, and collaborate across teams to maintain uptime and performance.
Top Skills: C#Ci/CdCloudGoJavaMicrosoft DefenderPython
Reposted 2 Days AgoSaved
In-Office
Boston, MA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
The Senior Site Reliability Engineer ensures the reliability and performance of cloud-native Kubernetes platforms by building tools, facilitating self-service for engineers, and promoting best practices.
Top Skills: ArgocdAWSAzureC#Ci/CdGitGoJavaKubernetesPulumiPythonTerraform
Reposted 2 Days AgoSaved
In-Office
Boston, MA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native Kubernetes platforms and tooling to ensure reliability, performance, and security. Write maintainable code, use CI/CD and IaC, employ observability for debugging, document systems, influence engineering practices, and improve platform reliability and cost efficiency.
Top Skills: AksApmAWSAzureC#Ci/CdCloud-NativeContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Reposted 2 Days AgoSaved
In-Office
Seattle, WA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native, production-grade Kubernetes platforms and tooling to improve reliability, operability, and developer experience. Develop IaC and CI/CD automation, use observability to debug distributed systems, document self-service workflows, and influence engineering teams on scalable architectural patterns. Scope and deliver platform projects with focus on security, cost efficiency, and service stability.
Top Skills: AksApmAWSAzureC#Ci/CdContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Reposted 2 Days AgoSaved
In-Office
Seattle, WA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
The Senior Site Reliability Engineer I will enhance Axon's observability platform, work on distributed tracing, log aggregation, and metrics infrastructure, and develop internal tools while collaborating with engineering teams.
Top Skills: ArgocdCdkCortexGoGrafanaHelmJaegerJavaLokiOpentelemetryPrometheusPythonTerraform
Reposted 2 Days AgoSaved
In-Office or Remote
3 Locations
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Owns reliability and operations for Azure Resource Manager and large-scale distributed cloud services. Responsibilities include on-call incident response, monitoring, observability, automation, performance optimization, service lifecycle ownership, coding, design documentation, and cross-functional delivery. The role requires improving availability, security, efficiency, and customer support across cloud environments while guiding engineers and maintaining service parity with the commercial cloud.
Top Skills: AzureAzure Resource ManagerJavaScriptPythonService FabricShell Scripting
3 Days AgoSaved
In-Office or Remote
Remote, OR, USA
Senior level
Senior level
Software
Define and operationalize reliability for a GPU-accelerated AI platform by creating SLIs, SLOs, error-budget practices, and high-quality alerting. Build APIs exposing reliability state to platform administrators, partner with infrastructure, storage, and networking teams on telemetry, and diagnose observability and performance issues across Kubernetes, bare-metal, and NVIDIA infrastructure in hybrid, edge, and air-gapped environments.
Top Skills: Cluster ApiGoGrafanaInfinibandK0Rdent AiK0Rdent EnterpriseK0Rdent Observability FrameworkKubernetesNvidia BmcNvlinkOpentelemetryPrometheusPythonRedfishUfmVictorialogsVictoriametrics
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 25 Days AgoSaved
In-Office
Charlotte, NC, USA
Senior level
Senior level
Insurance
Lead SRE role owning enterprise observability and monitoring strategy across cloud and applications. Architect Dynatrace-to-ServiceNow pipelines, alert correlation, paging integration, CMDB/service mapping, and automation for self-healing. Drive SRE practices (SLIs/SLOs/error budgets), mentor teams, reduce alert noise, improve MTTR, and standardize monitoring and incident workflows.
Top Skills: ArmBicepCmdbDynatraceGithub ActionsAzureMicrosoft TeamsPagerdutyRunbooksServicenow Itsm/ItomTerraform
Reposted 25 Days AgoSaved
Hybrid
Santa Clara, CA, USA
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software
Contract SRE to build, operate, and automate colo, on‑prem GPU clusters, and cloud infrastructure. Own provisioning, IaC (Terraform/Ansible), observability (Prometheus/Grafana or DataDog), incident response, runbooks, and customer-facing platform reliability.
Top Skills: AnsibleBashDatadogGrafanaKubernetesLinuxPrometheusPythonSystemdTerraform
Reposted 25 Days AgoSaved
In-Office
Santa Clara, CA, USA
230K-250K Annually
Senior level
230K-250K Annually
Senior level
Software • Analytics
Build and lead the SRE function: define SLOs/SLIs/error budgets, own observability and incident response, embed reliability in SDLC, and scale the SRE team for a distributed SaaS platform.
Top Skills: AnsibleAWSAzureBashDatadogDnsFirewallsGCPGrafanaKubernetesLoad BalancingPrometheusPythonRoutingSplunkSwitchingTcp/IpTerraform
Reposted 26 Days AgoSaved
In-Office
New York, NY, USA
Senior level
Senior level
Cloud • Information Technology • Consulting • Cybersecurity
Design, templatize and deploy scalable infrastructure in public clouds (AWS, GCP) using IaC (CloudFormation). Support architects, troubleshoot developer escalations, ensure compliance, and build stable platform services; work within agile teams to create configuration templates and automated deployments.
Top Skills: AWSAws CloudformationAws EfsEc2GCPPythonRdsRuby
Reposted 26 Days AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills: CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Reposted 26 Days AgoSaved
Remote
USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills: AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
Reposted 26 Days AgoSaved
Remote
USA
110K-140K Annually
Senior level
110K-140K Annually
Senior level
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills: AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
Reposted 26 Days AgoSaved
In-Office
Santa Clara, CA, USA
180K-260K Annually
Senior level
180K-260K Annually
Senior level
Automotive • Logistics
Operate, monitor, and scale on-prem and cloud infrastructure for an autonomous vehicle fleet. Manage data offload systems, CI environments, deployments, ETL/BI pipelines, dashboards, and performance optimization to ensure reliability, security, and operational efficiency.
Top Skills: AirflowArgoBashCiDockerDocker ComposeETLGrafanaHelmInfluxdbKanikoKubernetesPostgresPythonTimescaledb
Reposted 26 Days AgoSaved
In-Office
Hawthorne, CA, USA
125K-175K Annually
Junior
125K-175K Annually
Junior
Aerospace • Other
Design, operate, and scale on-premise infrastructure for the Starshield satellite constellation. Build automation for Kubernetes cluster deployment and management, operate core infrastructure (databases, monitoring, distributed storage), collaborate with software teams, troubleshoot across the stack, improve service lifecycle, and ensure high availability through monitoring and performance improvements.
Top Skills: AnsibleBashC++GoKubernetesLinuxOci ContainersPythonTerraform
Reposted 26 Days AgoSaved
In-Office
Redmond, WA, USA
125K-175K Annually
Junior
125K-175K Annually
Junior
Aerospace • Other
Design, deploy, and operate on-premises Kubernetes clusters and core infrastructure (databases, monitoring, distributed storage). Build automation, troubleshoot across the Starshield stack, collaborate with software teams to ensure scalable, highly available services, and improve lifecycle processes.
Top Skills: AnsibleBashBazelC++DatabasesDistributed StorageGoKubernetesLinuxMakefilesMonitoringOci ContainersPythonTcp/IpTerraform
Reposted 26 Days AgoSaved
Hybrid
New York, NY, USA
Senior level
Senior level
Fintech
The SRE/DevOps Engineer will enhance observability and monitoring tools, improve system reliability, conduct post-incident reviews, and collaborate with developers to optimize workflows and CI/CD processes.
Top Skills: AWSAzureAzure BicepAzure DevopsChaos MeshCloud FormationDatadogDockerElasticsearchGCPGithub ActionsGitlab Ci/CdGrafanaGremlinJenkinsKafkaKubernetesTerraform
Reposted 26 Days AgoSaved
In-Office or Remote
2 Locations
80K-133K Annually
Mid level
80K-133K Annually
Mid level
Consulting
Maintain and improve reliability of cloud-based enterprise systems by implementing SRE practices. Participate in design and code reviews, incident management, automation (IaC/CI-CD), monitoring, documentation, and collaboration with cross-functional teams to reduce downtime and improve scalability and security.
Top Skills: Ansible Automation PlatformArtifactoryAWSAzureBashCi/CdGitlabIacLinuxPackerPowershellPythonTerraformWindows
Reposted 26 Days AgoSaved
In-Office
Fairfax, VA, USA
Senior level
Senior level
Information Technology
Lead rearchitecture of a large monolith into cloud-native microservices. Design scalable, highly available systems on AWS and Azure using Kubernetes (AKS/EKS), Terraform, and robust cloud networking. Implement monitoring, logging, and DevOps practices, ensure container and cluster security, evaluate technologies, and mentor engineering teams.
Top Skills: AksAWSAzureCloud NetworkingEksKubernetesLoad BalancingTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account