Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Big Data • Analytics • Business Intelligence • Big Data Analytics
Seeking a Site Reliability Engineer to manage AI platform reliability, automate tasks, optimize ML pipelines, and lead incident response in a hybrid engineering role.
Top Skills:
ArgocdBigQueryCloud BuildDockerDvcGithub ActionsGoGrafanaKubeflowKubernetesMlflowPrometheusPub/SubPythonTerraformVertex Ai
Other
The Sr. Site Reliability Engineer will maintain and administer enterprise systems, troubleshoot operational issues, and develop scripts. This role requires collaboration across teams and participation in project planning and execution.
Top Skills:
AnsibleApacheAzureC#ChefIisJavaJbossPerlPowershellPuppetPythonRubyTomcat
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Own reliability, observability, and operational excellence for the Unified Call (911) platform. Design monitoring, alerting, incident response, deployment automation, and dashboards. Improve system resiliency, analyze architecture for operational risks, and build tooling and practices to enable reliable production operations across a Kubernetes-based cloud environment.
Top Skills:
AWSDatadogKafkaKubernetesRabbitMQ
Cloud • Fintech • HR Tech
Senior SRE for Workday Government Prism Analytics: build and maintain a secure, highly-available Kubernetes platform; automate infrastructure (Terraform, Argo CD); develop CI/CD pipelines; implement monitoring, alerting, and tracing; troubleshoot production issues; participate in on-call duties; ensure security/compliance for GovCloud deployments and collaborate across teams.
Top Skills:
SparkArgo CdAWSCi/CdDockerGoGrafanaHadoopHdfsKubernetesObservabilityOrmPrometheusPythonStormTerraform
Cloud • Fintech • HR Tech
Senior SRE for federal Workday Prism Analytics: design, build, and automate a secure, highly available Kubernetes platform; develop CI/CD pipelines and infrastructure-as-code; improve monitoring, alerting, and tracing; support production on-call and troubleshooting; collaborate across teams to maintain compliance and deliver analytics to GovCloud.
Top Skills:
SparkArgo CdAWSCi/CdDockerGoGrafanaKubernetesObservabilityPrometheusPythonTerraformTracing
Cloud • Fintech • HR Tech
Responsible for designing, building, automating, and maintaining a secure, highly available Kubernetes-based analytics platform for federal deployments. Improve CI/CD, observability, provisioning (Terraform, Argo CD), troubleshooting, on-call incident response, and collaborate across teams to deliver Workday Prism Analytics in GovCloud.
Top Skills:
SparkArgo CdAWSCi/CdDockerGoGovcloudGrafanaKubernetesObservabilityPrism AnalyticsPrometheusPythonTerraformTracing
Reposted 17 Days AgoSaved
Robotics • Software
Own reliability across vehicle and cloud stacks for AUV operations: onboard Jetson/ROS2 compute, topside systems, cloud ingestion/processing and customer platform. Build automation, observability, runbooks, and self-recovery to reduce on-call toil; manage AWS infrastructure, IaC, container orchestration, and reliability targets. Participate in shared 12-hour on-call shifts and field deployments, mentor team on operational excellence.
Top Skills:
AWSBashContainerizationDockerGoGrafanaIamJetsonKubernetesLinuxPrometheusPythonRosRos 2Terraform
Healthtech
As a Senior Site Reliability Engineer, you will ensure the reliability and performance of our Azure-based healthcare platform, implementing SRE practices, driving incident management, and automating operational tasks.
Top Skills:
AzureAzure MonitorBashDatadogPowershellPythonTerraform
Payments
Senior SRE responsible for ensuring high availability and resiliency of a global payments platform by building observability, automations, AI-driven remediation, incident response, and self-healing workflows; participates in on-call rotation and hybrid Philadelphia-based work.
Top Skills:
AiopsAksAnthropic (Claude)ApmAzureAzure Ai (Foundry)Azure Sre AgentCi/CdDatadogDnsDynatraceHTTPHttpsIisKubernetesLoad BalancingNew RelicOpenai (Codex)Pagerduty Process AutomationPowershellPythonRundeckSQLT-SqlTcp/IpVMwareWindows Server
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
Senior SRE responsible for designing and maturing cloud reliability on GCP/Azure: build observability, define SLIs/SLOs, create Terraform modules and CI/CD automation, lead incident/root-cause investigations, partner with security/governance, mentor engineers, and drive platform resiliency and production readiness.
Top Skills:
Azure Log AnalyticsAzure Resource GraphCi/CdDevsecopsDnsDynatraceFirewallsGenaiGoogle Cloud Platform (Gcp)IamLoad BalancingAzurePolicy-As-CodeTerraformTerraform EnterpriseVpc
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills:
AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Big Data • Software
Own and improve infrastructure for the Data Replication platform: Kubernetes, CI/CD, secrets, networking, cloud (AWS/GCP). Drive reliability, observability, AI-augmented tooling, canary rollouts, incident reduction, runbooks, and partner with product engineers.
Top Skills:
Agentic FrameworksAirbyteAWSCdksCi/CdConnector-Based ArchitecturesDatadogGCPGrafanaHelmJavaKubernetesLlmsPrometheusPythonSecrets ManagementTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native Kubernetes platforms and tooling to ensure reliability, performance, and security. Write maintainable code, use CI/CD and IaC, employ observability for debugging, document systems, influence engineering practices, and improve platform reliability and cost efficiency.
Top Skills:
AksApmAWSAzureC#Ci/CdCloud-NativeContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Build and operate cloud-native, production-grade Kubernetes platforms and tooling to improve reliability, operability, and developer experience. Develop IaC and CI/CD automation, use observability to debug distributed systems, document self-service workflows, and influence engineering teams on scalable architectural patterns. Scope and deliver platform projects with focus on security, cost efficiency, and service stability.
Top Skills:
AksApmAWSAzureC#Ci/CdContainer OrchestrationEksGoInfrastructure As CodeJavaKubernetesLoggingMetricsPulumiPythonTerraform
Information Technology • Legal Tech • Analytics
Design, build, and operate highly available AWS systems. Write and maintain Terraform, improve observability (Grafana, Pingdom, Uptrends), run on-call incident response, define SLOs/SLIs, build CI/CD with Azure DevOps/GitHub, automate operational work, document in Confluence, and mentor engineers.
Top Skills:
AWSAzure DevopsCi/CdConfluenceDockerGitGitGrafanaJIRAKubernetesLinuxPingdomServicenowTerraformUptrends
Artificial Intelligence • Healthtech
Own production reliability and performance for a healthcare AI platform: define SLIs/SLOs, run incident response and postmortems, build observability and automation, optimize scalability across serverless and container services, and partner with engineering, security, and compliance to improve operational maturity.
Top Skills:
Amazon EcsAurora PostgresAws LambdaBashClickhouseCloudwatchDatadogGithub ActionsGrafanaOpentelemetryPostgresPythonSentrySIEMVanta
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills:
DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Lead reliability and implementation for Kong's Managed Gateways: design and operate multi-cloud, Kubernetes-based systems, own incident response and SLOs, automate CI/CD and IaC, mentor SREs, and drive enterprise customer onboarding and technical implementations.
Top Skills:
AnsibleAWSAzureCassandraCi/CdDatadogElkGCPGoGrafanaIstioKubernetesLinkerdPostgresPrometheusTerraform
Defense • Manufacturing
Design, deploy, and maintain cloud and on-prem infrastructure with IaC; scale and operate mission-critical services; automate tooling and self-service platforms; collaborate with engineering teams to ensure high availability, reliability, security, and resiliency for vehicle and non-vehicle software systems.
Top Skills:
AWSAzureContainer OrchestrationGCPInfrastructure-As-Code (Iac)KubernetesLinuxNetworking
Financial Services
The Senior Site Reliability Engineer will enhance production insights, manage scalable infrastructure, optimize Kubernetes, and develop automation tools while ensuring high availability and performance in cloud-based systems.
Top Skills:
AnsibleAWSGCPGitopsGoGrafanaHelmIacKubernetesPythonSplunkTerraformTerragrunt
Aerospace • Other
Build, operate, and scale mission-critical application platforms to accelerate vehicle software delivery. Manage infrastructure as code, improve observability, collaborate with developers, run on-call rotations, conduct blameless postmortems, and reduce performance bottlenecks to support Falcon, Starship, Dragon, and Starlink software lifecycles.
Top Skills:
AnsibleBazelBuckC#C++ClickhouseDockerJavaScriptKubernetesKvmLinuxMakeMySQLPostgresPuppetPythonQemuTerraformVsphere
Reposted 20 Days AgoSaved
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills:
AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills:
AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Artificial Intelligence • Cloud • Social Impact • Software • Wearables
Lead design and implementation of a Kubernetes-based self-service platform, driving infrastructure-as-code, GitOps practices, datastore provisioning, observability, and architectural standards. Partner with application teams to resolve performance bottlenecks and build scalable, automated platform solutions for large-scale distributed systems.
Top Skills:
AWSAzureCassandraDatadogGCPGitopsGrafanaKafkaKubernetesMicroservicesMySQLNew RelicPostgresPulumiTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results



















.png)











