Top Site Reliability Engineer Jobs

Reposted 16 Days AgoSaved
In-Office
Washington, DC, USA
Senior level
Senior level
Big Data • Analytics • Business Intelligence • Big Data Analytics
Seeking a Site Reliability Engineer to manage AI platform reliability, automate tasks, optimize ML pipelines, and lead incident response in a hybrid engineering role.
Top Skills: ArgocdBigQueryCloud BuildDockerDvcGithub ActionsGoGrafanaKubeflowKubernetesMlflowPrometheusPub/SubPythonTerraformVertex Ai
Reposted 16 Days AgoSaved
In-Office
Seattle, WA, USA
Senior level
Senior level
Other
The Sr. Site Reliability Engineer will maintain and administer enterprise systems, troubleshoot operational issues, and develop scripts. This role requires collaboration across teams and participation in project planning and execution.
Top Skills: AnsibleApacheAzureC#ChefIisJavaJbossPerlPowershellPuppetPythonRubyTomcat
Reposted 16 Days AgoSaved
In-Office
New York, NY, USA
141K-217K Annually
Senior level
141K-217K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Own reliability, observability, and operational excellence for the Unified Call (911) platform. Design monitoring, alerting, incident response, deployment automation, and dashboards. Improve system resiliency, analyze architecture for operational risks, and build tooling and practices to enable reliable production operations across a Kubernetes-based cloud environment.
Top Skills: AWSDatadogKafkaKubernetesRabbitMQ
17 Days AgoSaved
In-Office
Reston, VA, USA
137K-244K Annually
Senior level
137K-244K Annually
Senior level
Cloud • Fintech • HR Tech
Senior SRE for Workday Government Prism Analytics: build and maintain a secure, highly-available Kubernetes platform; automate infrastructure (Terraform, Argo CD); develop CI/CD pipelines; implement monitoring, alerting, and tracing; troubleshoot production issues; participate in on-call duties; ensure security/compliance for GovCloud deployments and collaborate across teams.
Top Skills: SparkArgo CdAWSCi/CdDockerGoGrafanaHadoopHdfsKubernetesObservabilityOrmPrometheusPythonStormTerraform
17 Days AgoSaved
In-Office
Reston, VA, USA
137K-244K Annually
Senior level
137K-244K Annually
Senior level
Cloud • Fintech • HR Tech
Senior SRE for federal Workday Prism Analytics: design, build, and automate a secure, highly available Kubernetes platform; develop CI/CD pipelines and infrastructure-as-code; improve monitoring, alerting, and tracing; support production on-call and troubleshooting; collaborate across teams to maintain compliance and deliver analytics to GovCloud.
Top Skills: SparkArgo CdAWSCi/CdDockerGoGrafanaKubernetesObservabilityPrometheusPythonTerraformTracing
17 Days AgoSaved
In-Office
Reston, VA, USA
137K-244K Annually
Senior level
137K-244K Annually
Senior level
Cloud • Fintech • HR Tech
Responsible for designing, building, automating, and maintaining a secure, highly available Kubernetes-based analytics platform for federal deployments. Improve CI/CD, observability, provisioning (Terraform, Argo CD), troubleshooting, on-call incident response, and collaborate across teams to deliver Workday Prism Analytics in GovCloud.
Top Skills: SparkArgo CdAWSCi/CdDockerGoGovcloudGrafanaKubernetesObservabilityPrism AnalyticsPrometheusPythonTerraformTracing
Reposted 17 Days AgoSaved
Remote
USA
164K-220K Annually
Senior level
164K-220K Annually
Senior level
Robotics • Software
Own reliability across vehicle and cloud stacks for AUV operations: onboard Jetson/ROS2 compute, topside systems, cloud ingestion/processing and customer platform. Build automation, observability, runbooks, and self-recovery to reduce on-call toil; manage AWS infrastructure, IaC, container orchestration, and reliability targets. Participate in shared 12-hour on-call shifts and field deployments, mentor team on operational excellence.
Top Skills: AWSBashContainerizationDockerGoGrafanaIamJetsonKubernetesLinuxPrometheusPythonRosRos 2Terraform
Reposted 17 Days AgoSaved
Hybrid
2 Locations
Senior level
Senior level
Healthtech
As a Senior Site Reliability Engineer, you will ensure the reliability and performance of our Azure-based healthcare platform, implementing SRE practices, driving incident management, and automating operational tasks.
Top Skills: AzureAzure MonitorBashDatadogPowershellPythonTerraform
Reposted 17 Days AgoSaved
Remote or Hybrid
Location, WV, USA
Senior level
Senior level
Payments
Senior SRE responsible for ensuring high availability and resiliency of a global payments platform by building observability, automations, AI-driven remediation, incident response, and self-healing workflows; participates in on-call rotation and hybrid Philadelphia-based work.
Top Skills: AiopsAksAnthropic (Claude)ApmAzureAzure Ai (Foundry)Azure Sre AgentCi/CdDatadogDnsDynatraceHTTPHttpsIisKubernetesLoad BalancingNew RelicOpenai (Codex)Pagerduty Process AutomationPowershellPythonRundeckSQLT-SqlTcp/IpVMwareWindows Server
Reposted 17 Days AgoSaved
In-Office
3 Locations
153K-192K Annually
Senior level
153K-192K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
Senior SRE responsible for designing and maturing cloud reliability on GCP/Azure: build observability, define SLIs/SLOs, create Terraform modules and CI/CD automation, lead incident/root-cause investigations, partner with security/governance, mentor engineers, and drive platform resiliency and production readiness.
Top Skills: Azure Log AnalyticsAzure Resource GraphCi/CdDevsecopsDnsDynatraceFirewallsGenaiGoogle Cloud Platform (Gcp)IamLoad BalancingAzurePolicy-As-CodeTerraformTerraform EnterpriseVpc
Reposted 18 Days AgoSaved
In-Office or Remote
8 Locations
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills: AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Reposted 18 Days AgoSaved
In-Office or Remote
2 Locations
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills: AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 18 Days AgoSaved
Hybrid
San Francisco, CA, USA
196K-235K Annually
Senior level
196K-235K Annually
Senior level
Artificial Intelligence • Big Data • Software
Own and improve infrastructure for the Data Replication platform: Kubernetes, CI/CD, secrets, networking, cloud (AWS/GCP). Drive reliability, observability, AI-augmented tooling, canary rollouts, incident reduction, runbooks, and partner with product engineers.
Top Skills: Agentic FrameworksAirbyteAWSCdksCi/CdConnector-Based ArchitecturesDatadogGCPGrafanaHelmJavaKubernetesLlmsPrometheusPythonSecrets ManagementTerraform
19 Days AgoSaved
In-Office
Boston, MA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native Kubernetes platforms and tooling to ensure reliability, performance, and security. Write maintainable code, use CI/CD and IaC, employ observability for debugging, document systems, influence engineering practices, and improve platform reliability and cost efficiency.
Top Skills: AksApmAWSAzureC#Ci/CdCloud-NativeContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
19 Days AgoSaved
In-Office
Seattle, WA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native, production-grade Kubernetes platforms and tooling to improve reliability, operability, and developer experience. Develop IaC and CI/CD automation, use observability to debug distributed systems, document self-service workflows, and influence engineering teams on scalable architectural patterns. Scope and deliver platform projects with focus on security, cost efficiency, and service stability.
Top Skills: AksApmAWSAzureC#Ci/CdContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Reposted 19 Days AgoSaved
In-Office or Remote
15 Locations
100K-210K Annually
Senior level
100K-210K Annually
Senior level
Information Technology • Legal Tech • Analytics
Design, build, and operate highly available AWS systems. Write and maintain Terraform, improve observability (Grafana, Pingdom, Uptrends), run on-call incident response, define SLOs/SLIs, build CI/CD with Azure DevOps/GitHub, automate operational work, document in Confluence, and mentor engineers.
Top Skills: AWSAzure DevopsCi/CdConfluenceDockerGitGitGrafanaJIRAKubernetesLinuxPingdomServicenowTerraformUptrends
19 Days AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Healthtech
Own production reliability and performance for a healthcare AI platform: define SLIs/SLOs, run incident response and postmortems, build observability and automation, optimize scalability across serverless and container services, and partner with engineering, security, and compliance to improve operational maturity.
Top Skills: Amazon EcsAurora PostgresAws LambdaBashClickhouseCloudwatchDatadogGithub ActionsGrafanaOpentelemetryPostgresPythonSentrySIEMVanta
Reposted 19 Days AgoSaved
Remote
2 Locations
150K-195K Annually
Senior level
150K-195K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills: DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
20 Days AgoSaved
In-Office or Remote
Washington, DC, USA
113K-162K Annually
Senior level
113K-162K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Lead reliability and implementation for Kong's Managed Gateways: design and operate multi-cloud, Kubernetes-based systems, own incident response and SLOs, automate CI/CD and IaC, mentor SREs, and drive enterprise customer onboarding and technical implementations.
Top Skills: AnsibleAWSAzureCassandraCi/CdDatadogElkGCPGoGrafanaIstioKubernetesLinkerdPostgresPrometheusTerraform
20 Days AgoSaved
In-Office
Los Angeles, CA, USA
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Defense • Manufacturing
Design, deploy, and maintain cloud and on-prem infrastructure with IaC; scale and operate mission-critical services; automate tooling and self-service platforms; collaborate with engineering teams to ensure high availability, reliability, security, and resiliency for vehicle and non-vehicle software systems.
Top Skills: AWSAzureContainer OrchestrationGCPInfrastructure-As-Code (Iac)KubernetesLinuxNetworking
Reposted 20 Days AgoSaved
In-Office
New York, NY, USA
140K-170K Annually
Senior level
140K-170K Annually
Senior level
Financial Services
The Senior Site Reliability Engineer will enhance production insights, manage scalable infrastructure, optimize Kubernetes, and develop automation tools while ensuring high availability and performance in cloud-based systems.
Top Skills: AnsibleAWSGCPGitopsGoGrafanaHelmIacKubernetesPythonSplunkTerraformTerragrunt
Reposted 20 Days AgoSaved
In-Office
Hawthorne, CA, USA
165K-230K Annually
Senior level
165K-230K Annually
Senior level
Aerospace • Other
Build, operate, and scale mission-critical application platforms to accelerate vehicle software delivery. Manage infrastructure as code, improve observability, collaborate with developers, run on-call rotations, conduct blameless postmortems, and reduce performance bottlenecks to support Falcon, Starship, Dragon, and Starlink software lifecycles.
Top Skills: AnsibleBazelBuckC#C++ClickhouseDockerJavaScriptKubernetesKvmLinuxMakeMySQLPostgresPuppetPythonQemuTerraformVsphere
Reposted 20 Days AgoSaved
Remote
USA
117K-181K Annually
Senior level
117K-181K Annually
Senior level
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills: AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Reposted 20 Days AgoSaved
In-Office or Remote
7 Locations
Senior level
Senior level
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills: AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Reposted 21 Days AgoSaved
In-Office
Seattle, WA, USA
134K-215K Annually
Senior level
134K-215K Annually
Senior level
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Lead design and implementation of a Kubernetes-based self-service platform, driving infrastructure-as-code, GitOps practices, datastore provisioning, observability, and architectural standards. Partner with application teams to resolve performance bottlenecks and build scalable, automated platform solutions for large-scale distributed systems.
Top Skills: AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesMicroservicesMySQLNew RelicPostgresPulumiTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account