Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Healthtech • Pet • Biotech
Senior SRE responsible for designing and modernizing CI/CD and deployment systems, automating AWS Serverless infrastructure, improving observability and incident response, enforcing release and security practices, and guiding engineering teams to scale resilient global services.
Top Skills:
AuroradbAws CloudformationAws LambdaAzure Entra IdCloudfrontDynamoDBEventbridgeGitGitGithub ActionsMavenOauth2Openid ConnectS3SnsSqsTerraform
Digital Media • Events • Music
Lead and manage a team of SRE/DevOps engineers to ensure reliability, availability, and performance of cloud-based systems. Oversee incident response, operational troubleshooting, process improvements, and cross-team collaboration while mentoring and delegating tasks to meet business objectives.
Top Skills:
Cloud Services
Software • Cybersecurity
Senior SRE responsible for ensuring reliability, scalability, and performance of cloud systems (AWS/Azure). Duties include monitoring, on-call support, incident response and RCA, releases and maintenance, security controls, and automation to improve infrastructure efficiency.
Top Skills:
AWSAzureAzure DevopsDockerGitlab Ci/CdGoJenkinsKubernetesPowershellPython
Software • Cybersecurity
Drive reliability, scalability, and performance of cloud-based systems on AWS/Azure. Monitor systems, handle on-call production support, lead incident response and root cause analysis, perform releases and hotfixes, implement cloud security controls, and automate infrastructure improvements.
Top Skills:
AWSAzureAzure DevopsCloud-NativeDockerGitlab Ci/CdGoJenkinsKubernetesMicroservicesPowershellPython
Cloud • Fintech • HR Tech
The Senior Site Reliability Engineer will maintain and optimize the Kubernetes-based platform, ensuring high availability, automating infrastructure, and adhering to security compliance, while troubleshooting and collaborating with development teams.
Top Skills:
Argo CdAWSC#Ci/CdGoKubernetesPythonRubyRustTerraform
Social Media
Operate, scale, and harden an AWS- and Kubernetes-based platform using GitOps. Build CI/CD and infrastructure-as-code (Terraform/Terragrunt), manage ArgoCD/Helm deployments, improve observability, automate toil reduction, lead incident response and post-incident remediation, and partner with application, security, and platform teams to improve reliability and delivery.
Top Skills:
ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonRbacTerraformTerragrunt
Cloud • Information Technology • Internet of Things • Software • Consulting • Infrastructure as a Service (IaaS) • Automation
Design, build, automate, and operate Red Hat Hybrid OpenShift platforms across cloud and on-prem. Implement GitOps, CI/CD, monitoring, and SRE practices; develop operators and tooling; lead incident response, on-call duties, and postmortems; mentor peers and improve platform reliability and self-service.
Top Skills:
ArgocdAWSAzureCatchpointCi/CdDatadogDnsFedoraGitlabGitopsGoGoGCPGrafanaHttp/TlsKubernetesKubernetes OperatorsLdapLinuxOpenshiftOpenshift PipelinesOperator SdkPrometheusPythonRhelSplunkSplunk ImTcp/IpTekton
Fintech • Software
Senior SRE responsible for ensuring reliability, scalability, and performance of production systems. Investigates and resolves incidents, works with R&D on defects, manages deployments and change validation, builds monitoring and diagnostic tooling to improve MTTA/MTTD/MTTR, configures observability platforms, participates in on-call rotation, and performs risk assessments and production readiness validation.
Top Skills:
Ai ToolsAkamaiAmqAnsibleAWSAzureBashCdnCloudflareCloudwatchDatadogDatastreamDnsDockerDynatraceEfsEksHTTPHttpsInterconnectJavaJmsKubernetesMongoDBOraclePostgresPrometheusPythonRabbitMQRedshiftS3SplunkTcpTerraformUdpUnix/LinuxWafZabbix
Fintech
Lead the reliability, monitoring, and scaling of Kubernetes clusters and cloud infrastructure. Deploy and maintain containerized microservices, implement IaC (Terraform/CloudFormation), automate operational tasks, manage incidents, and ensure security and compliance across environments. Collaborate with development and operations teams to optimize performance and availability.
Top Skills:
Amazon S3AnsibleApache MesosAWSAzureC/C++CephCloudFormationDockerGCPHdfsHelmJavaJavaScriptJenkinsKubernetesLinuxNfsPostgresPythonRubyTerraformYarn
Software
As a Site Reliability Engineer, you'll build and maintain infrastructure for ML models, automate processes, and collaborate cross-functionally.
Top Skills:
Circle CiCloudFormationElk StackGithub ActionsGitlab CiGrafanaJenkinsKubernetesOpentelemetryPrometheusPulumiTerraform
Healthtech • Software • Analytics • Business Intelligence
Senior SRE responsible for designing, building, and operating reliable, scalable distributed systems; owning production reliability (SLOs/SLIs, incident response, MTTR reduction); automating toil with software and platform tooling; driving observability, capacity planning, and cross-team reliability improvements; mentoring engineers and running blameless postmortems.
Top Skills:
AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Cloud • Software
Design, implement, and support Kubernetes and compute platforms in a private cloud. Oversee architecture and standardization across hardware, OS, and cloud orchestration.
Top Skills:
AnsibleBashCi/CdHelmKubernetesLinuxOpenstackPythonTerraformUbuntu
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
The Senior Site Reliability Engineer will design and maintain scalable infrastructure, improve system reliability, manage CI/CD pipelines, and collaborate across teams for operational excellence.
Top Skills:
AnsibleArgocdAWSBashDatadogDockerElkGithub ActionsGrafanaKubernetesLinuxOpentelemetryPrometheusPythonTerraform
Artificial Intelligence • Machine Learning • Natural Language Processing • Software
As a Site Reliability Engineer, you'll enhance the performance and reliability of infrastructure and products by collaborating with engineering teams, automating configurations, and implementing monitoring systems.
Top Skills:
AWSGoKubernetesPythonTerraform
Software • Cybersecurity
The Senior Site Reliability Engineer will ensure the reliability and performance of cloud systems, manage AWS/Azure infrastructure, optimize performance, lead incident responses, and implement security best practices.
Top Skills:
AWSAzureAzure DevopsDockerGitlab Ci/CdGoJenkinsKubernetesPowershellPython
Reposted 15 Days AgoSaved
Financial Services
Own reliability and scalability of on-prem observability platforms (ELK, Grafana); handle production escalations, capacity planning, SLOs, onboarding, automation, IaC (Terraform/Helm/Ansible), upgrades, security hardening, and platform modernization.
Top Skills:
AnsibleApm InstrumentationBashBeatsChefElasticsearchElk StackFluent BitFluentdGrafanaHelmKibanaLinuxLogstashNew RelicOpentelemetryPrometheusPuppetPythonShell Scripting/Linux ShellSolarwindsTerraform
Reposted 15 Days AgoSaved
Other • Social Impact
As a Senior Site Reliability Engineer, you will manage and improve Wikimedia's infrastructure, handle operational tasks, automate processes, and provide mentorship while participating in a 24/7 on-call rotation.
Top Skills:
AnsibleBashDebianGoGrafanaHhvmKubernetesMemcachedPHPPrometheusPuppetPythonRedisRuby
Big Data • Cloud • Marketing Tech • Social Impact • Software
The Senior Site Reliability Engineer will support global product deployments, provide 24/7 operational support, maintain CI/CD tooling, and optimize system performance. They will utilize their SRE practices and leadership abilities to improve product reliability and guide other engineers.
Top Skills:
AWSCassandraCircleCIDynamoDBGCPGoJenkinsKubernetesNosql DatabasesPythonScylladbSinglestore DbTerraform
Big Data • Cloud • Marketing Tech • Social Impact • Software
As a Senior Site Reliability Engineer, you will manage global product deployments, provide operational support, enhance CI/CD tooling, and optimize system performance, collaborating closely with distributed engineering teams.
Top Skills:
AWSCassandraCircleCIDynamoDBGCPGoJenkinsKubernetesNosql DatabasesPythonScylladbSinglestore DbTerraform
Big Data • Cloud • Hardware • Software • App development
Senior Site Reliability Engineer responsible for improving reliability and performance of wwt.com. Build monitoring/alerting, lead incident response and root cause analysis, automate operational tasks, run load testing, review penetration test results, mentor engineers, and collaborate with developers to ensure scalable, resilient systems.
Top Skills:
AngularGoJavaNode.jsPythonReactVue
Cloud • Information Technology • Productivity • Security • Software
Senior Site Reliability Engineer responsible for improving reliability, performance, and scalability of wwt.com. Develop monitoring, alerting, and automation; lead incident response and root cause analysis; perform load testing and security vulnerability remediation; collaborate with developers and QA; write code to automate operations; mentor engineers and influence technical direction.
Top Skills:
AngularGoJavaNode.jsPythonReactVue
Aerospace • Other
Design, deploy, and operate highly available, sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve performance, and advance deployment, monitoring, and alerting. Collaborate across teams through full software lifecycle to build scalable, operable services for Starlink.
Top Skills:
AlertingApache FlinkApache KafkaSparkBare Metal Compute ClustersC#Continuous IntegrationGoHbaseHdfsIstioJavaKubernetesLinuxMonitoringPythonScalaVersion Control
Aerospace • Other
Design, deploy, and operate sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve deployment/monitoring/alerting, collaborate across teams, and optimize performance throughout the software lifecycle.
Top Skills:
Apache KafkaSparkBare MetalC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Aerospace • Other
Design, upgrade, and operate large-scale distributed systems for Starlink. Improve sharding, geo-redundancy, multi-region deployment, monitoring, and performance. Manage petabyte-scale bare-metal clusters and collaborate across teams through the full software lifecycle.
Top Skills:
Apache KafkaSparkC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Software
Lead the modernization of AWS cloud infrastructure, implement automation, ensure system reliability, and manage performance with a focus on security and incident response.
Top Skills:
AngularApexAWSC#ElasticacheNew RelicNode.jsNpmPm2PythonRedisShell ScriptingTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results





























