Maximum of 25 job preferences reached.
Top Site Reliability Engineer Jobs
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, improve workflows, manage secure infrastructure, and participate in on-call rotation for an AI-driven company.
Top Skills:
AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIpIscsiJenkinsKubernetesLinux/DebianMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRaidRubyS3ScyllaSshSslSupermicroTcpTlsUbuntu
Big Data • Healthtech • Information Technology • Analytics
Serve as an SRE-focused AI enablement partner: evaluate AI architectures, coach engineering teams on LLM integrations and agentic/RAG patterns, advise on observability, SLOs, governance, and operational readiness, and develop internal standards and documentation.
Top Skills:
Agentic FrameworksAnthropic ClaudeAWSAzureAzure Ai FoundryCi/CdDatabricksDatadogDockerGrafanaKubernetesLangchainLlamaindexLlm ApiOpentelemetryRagSemantic Kernel
Cloud
Build and operate reliable, scalable, secure infrastructure for SaaS security and Snowflake data systems. Automate infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and related tools. Ensure security and compliance, participate in on-call rotations, lead incident response and root-cause analysis, and collaborate with development, data science, and security teams on architectural decisions and service implementation.
Top Skills:
Artificial IntelligenceCi/CdFlywayInfrastructure As CodeKubernetesMachine LearningSnowflakeSpinnakerTerraform
Cloud
Designs, builds, and operates reliable, scalable infrastructure for security SaaS and Snowflake data systems. Automates infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and Flyway. Partners with security, development, and data science teams on secure, compliant architecture. Participates in on-call rotations, leads critical incident response and root-cause analysis, and implements preventative improvements. In-person onboarding and travel to the Toronto office are required during the first employment week.
Top Skills:
Ci/CdContainerizationFlywayInfrastructure As CodeKubernetesSnowflakeSpinnakerTerraform
Information Technology • Security • Cybersecurity
Own production operations and highly available cloud infrastructure for regulated government environments. Lead incident response, root-cause analysis, disaster recovery testing, compliance operationalization, audit readiness, vulnerability management, and continuous monitoring. Build automation, secure CI/CD pipelines, infrastructure-as-code, observability, and compliance tooling across Kubernetes, Linux, containers, and cloud platforms. Partner with security, compliance, and engineering teams to improve reliability, deployment safety, and regulatory sustainment.
Top Skills:
Aws GovcloudBashCi/CdDod Il4Dod Il5FedrampGitopsGoGrafanaKubernetesLinuxNist 800-53PrometheusPythonStigTerraformUnixZero Trust
Social Media
Operate, scale, and harden an AWS- and Kubernetes-based platform using GitOps. Build CI/CD and infrastructure-as-code (Terraform/Terragrunt), manage ArgoCD/Helm deployments, improve observability, automate toil reduction, lead incident response and post-incident remediation, and partner with application, security, and platform teams to improve reliability and delivery.
Top Skills:
ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonRbacTerraformTerragrunt
Cloud • Information Technology • Internet of Things • Software • Consulting • Infrastructure as a Service (IaaS) • Automation
Design, build, automate, and operate Red Hat Hybrid OpenShift platforms across cloud and on-prem. Implement GitOps, CI/CD, monitoring, and SRE practices; develop operators and tooling; lead incident response, on-call duties, and postmortems; mentor peers and improve platform reliability and self-service.
Top Skills:
ArgocdAWSAzureCatchpointCi/CdDatadogDnsFedoraGitlabGitopsGoGoGCPGrafanaHttp/TlsKubernetesKubernetes OperatorsLdapLinuxOpenshiftOpenshift PipelinesOperator SdkPrometheusPythonRhelSplunkSplunk ImTcp/IpTekton
Reposted One Month AgoSaved
Automotive • Cloud • Hardware • Software
Design, build, and operate developer platform infrastructure to support firmware build pipelines. Implement IaC patterns, GitOps, Kubernetes operators, CI optimization, cloud reliability, security policy-as-code, and developer tooling while mentoring engineers and collaborating on architecture and risk mitigation.
Top Skills:
ArgocdAWSAzureBashCiFluxGCPGitopsGoKubernetesPythonService MeshTerraform
Information Technology • Automation
The SRE/Infrastructure Engineer will architect and manage secure, scalable systems for automated penetration testing, optimizing reliability, and enhancing infrastructure based on customer demand. Responsibilities include maintaining production environments, leading technical discussions, and promoting high coding standards.
Top Skills:
AWSAzureCloudFormationElkGCPNew RelicOpentelemetryPostgresPrometheusTerraform
HR Tech • Information Technology • Professional Services • Software • Business Intelligence • Consulting • Automation
Seeking a Site Reliability Engineer with expertise in Unix/Linux, scripting languages, and experience in containerization, cloud platforms, and application monitoring tools.
Top Skills:
AnsibleApache TomcatAWSCassandraChefCoradiantDockerDynatraceElasticGomezGCPJenkinsKafkaLinuxMq SeriesOraclePuppetPythonShell ScriptingSplunkTealeafUnixVagrantWebsphere
Software
Operate and improve reliability across Retool Cloud, managed, BYOC, and self-hosted deployments. Automate provisioning, upgrades, migrations, and secret rotations. Build observability and safer deployment/rollback workflows, partner with product teams, and produce runbooks, docs, and migration guides to reduce customer toil and scale operations.
Top Skills:
AWSDocker ComposeGoHelmJavaKubernetesPostgresPythonRubyTerraformTypescript
Artificial Intelligence • Cloud • Information Technology • Software
Contribute to the reliability and performance of Mithril's GPU orchestration platform through automation, observability, and infrastructure management. Collaborate with the team to ensure scalability across multi-cloud environments while maintaining systems stability and implementing SLOs.
Top Skills:
AWSAzureGCPGoGrafanaKubernetesLinuxOpentelemetryPrometheusPulumiPythonTcp/IpTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Fintech • Financial Services
Support and operate OpenShift/Kubernetes platforms and RHEL systems for a trading exchange. Maintain automation (Ansible, Jenkins, ArgoCD, GitOps), cloud workloads (AWS/Azure), authentication/DNS/time services, enterprise hardware/storage, and developer platform integrations. Monitor, respond to incidents, perform patching and changes, and collaborate with developers, traders, and infrastructure teams to ensure security, resilience, and availability.
Top Skills:
AnsibleArgocdArtifactoryAWSAzureBashBindChronyCorednsDell PoweredgeGithub ActionsGithub EnterpriseGitopsJenkinsKerberosKubernetesLdapNtpPtpPythonRed Hat Enterprise Linux (Rhel)Red Hat OpenshiftSonarqubeSssdTerraform
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
Design, operate, and automate global network and reliability infrastructure for large-scale ML workloads and a private supercomputer. Own device configuration management, protocols (BGP, VPNs, WAN), datacenter fabrics, monitoring/SLOs, incident response, security/compliance, and cross-team reliability improvements.
Top Skills:
AirflowAnsibleBashBgpBluefieldCniCumulus LinuxDatadogEcmpElkEvpn/VxlanFirewallsGrafanaInfinibandInfobloxIngressIpsec VpnsIscsiKafkaKubernetesLinuxLoad BalancersLustrefsMplsNetboxNetwork PolicyNfsNornirOpentelemetryPrometheusPythonQosService NetworkingSparkSpectrum-XSpine-LeafSwitchesTerraformVpnsWan Circuits
Cloud
The Site Reliability Engineer at TCN will design, deploy, and maintain systems for performance, reliability, and security, while managing incidents and collaborating with teams.
Top Skills:
BashGoGoogle Cloud PlatformJavaKubernetesLinuxNode.jsPythonRuby
Fintech
The Site Reliability Engineer will monitor and manage Kubernetes clusters, optimize Cloud Infrastructure, and automate processes using tools like Terraform and Docker.
Top Skills:
Amazon S3AWSAzureC/C++CephDockerGCPHdfsHelmJavaJavaScriptKubernetesNfsPostgresPythonRubyTerraform
Information Technology • Cryptocurrency
The Site Reliability Engineer will lead technical initiatives, architect solutions, troubleshoot issues, mentor team members, and improve observability practices.
Top Skills:
ArgocdBashElk StackGCPGoGrafanaHelmKubernetesPrometheusPythonTerraform
Information Technology • Marketing Tech • Social Media
Lead technical design and architecture of large-scale Ceph clusters, drive capacity planning and major upgrades, resolve complex cross-functional production incidents, establish automation and reliability standards, and mentor engineers on advanced Ceph operations to scale the global storage platform.
Top Skills:
AnsibleBluestoreCephCephfsCinderCrushCsiDdnErasure CodingGoIsilonKubernetesManilaMdsMgrMonNetappNeutronNovaOpenstackOsdPlacement GroupsPurePythonRbdRbd MirroringRgwRgw MultisiteRookS3SaltstackSwiftTerraformVastWeka
Information Technology • Marketing Tech • Social Media
Own and maintain reliability, performance, scalability, and capacity of massive Ceph storage clusters. Troubleshoot distributed storage issues, build automation (Python/Shell, SaltStack, Ansible), define observability (SLIs/SLOs, PromQL/LogQL, Grafana), and lead lifecycle initiatives like upgrades, expansions, and migrations.
Top Skills:
AnsibleCeph (RadosCephfs)ChefGrafanaLinuxLogqlPromqlPuppetPythonRbdRgwSaltstackShell
Energy • Chemical • Utilities • Manufacturing
Design, implement, and maintain observability, alerting, and developer productivity systems for production and internal services. Instrument services with metrics, logs, and traces, run on-call, respond to incidents, lead reviews, and automate operational workflows to improve reliability and reduce toil.
Top Skills:
Ci/CdDatadogDnsGrafanaHTTPInfrastructure-As-CodeLoggingMetricsOpentelemetryPrometheusTlsTracing
Blockchain • Energy • Cryptocurrency
Hands-on role to assess, implement, test, and document backup, restore, failover, and recovery capabilities. Inventory critical systems, design and automate backup and restoration, run recovery exercises, produce runbooks, validate recoverability, measure RTO/RPO, and train system owners. Collaborate with Security, SRE, DevOps, QA, and application teams to harden shared recovery capabilities and transfer operational ownership.
Healthtech • Software • Analytics • Business Intelligence
Senior SRE responsible for designing, building, and operating reliable, scalable distributed systems; owning production reliability (SLOs/SLIs, incident response, MTTR reduction); automating toil with software and platform tooling; driving observability, capacity planning, and cross-team reliability improvements; mentoring engineers and running blameless postmortems.
Top Skills:
AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Fintech • Software
Senior SRE responsible for ensuring reliability, scalability, and performance of production systems. Investigates and resolves incidents, works with R&D on defects, manages deployments and change validation, builds monitoring and diagnostic tooling to improve MTTA/MTTD/MTTR, configures observability platforms, participates in on-call rotation, and performs risk assessments and production readiness validation.
Top Skills:
Ai ToolsAkamaiAmqAnsibleAWSAzureBashCdnCloudflareCloudwatchDatadogDatastreamDnsDockerDynatraceEfsEksHTTPHttpsInterconnectJavaJmsKubernetesMongoDBOraclePostgresPrometheusPythonRabbitMQRedshiftS3SplunkTcpTerraformUdpUnix/LinuxWafZabbix
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
The Senior Site Reliability Engineer will design and maintain scalable infrastructure, improve system reliability, manage CI/CD pipelines, and collaborate across teams for operational excellence.
Top Skills:
AnsibleArgocdAWSBashDatadogDockerElkGithub ActionsGrafanaKubernetesLinuxOpentelemetryPrometheusPythonTerraform
Fintech
Lead the reliability, monitoring, and scaling of Kubernetes clusters and cloud infrastructure. Deploy and maintain containerized microservices, implement IaC (Terraform/CloudFormation), automate operational tasks, manage incidents, and ensure security and compliance across environments. Collaborate with development and operations teams to optimize performance and availability.
Top Skills:
Amazon S3AnsibleApache MesosAWSAzureC/C++CephCloudFormationDockerGCPHdfsHelmJavaJavaScriptJenkinsKubernetesLinuxNfsPostgresPythonRubyTerraformYarn
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs
.NET Developer Jobs
Aerospace Thermal Engineering Jobs
AI Engineer Jobs
Android Developer Jobs
Automation Engineer Jobs
Backend Developer Jobs
Blockchain Developer Jobs
C# Jobs
C++ Jobs
Cloud Architect Jobs
Cloud Engineer Jobs
Design Engineer Jobs
DevOps Engineer Jobs
Director Of Engineering Jobs
Electrical Engineering Jobs
Embedded Software Engineer Jobs
Engineering Jobs
Engineering Manager Jobs
Environmental Engineering Jobs
Field Engineer Jobs
Front End Developer Jobs
Full Stack Developer Jobs
Game Developer Jobs
Golang Jobs
Hardware Engineer Jobs
Industrial Engineering Jobs
iOS Developer Jobs
Java Developer Jobs
Javascript Developer Jobs
Linux Jobs
Manufacturing Engineer Jobs
Mechanical Engineering Jobs
Network Engineer Jobs
PHP Developer Jobs
Process Engineer Jobs
Project Engineer Jobs
Prompt Engineering Jobs
Python Jobs
QA Jobs
Robotics Engineer Jobs
Ruby on Rails Jobs
Salesforce Administrator Jobs
Salesforce Developer Jobs
Scala Jobs
Sharepoint Developer Jobs
Site Reliability Engineer Jobs
Software Engineering Manager Jobs
Solutions Architect Jobs
SQL Developer Jobs
Structural Engineer Jobs
System Engineer Jobs
Test Engineer Jobs
Web Developer Jobs
All Filters
Total selected ()
No Results
No Results






.png)






.png)














