Top Site Reliability Engineer Jobs

Reposted 12 Days AgoSaved
In-Office
San Francisco, CA, USA
160K-250K Annually
Mid level
160K-250K Annually
Mid level
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, improve workflows, manage secure infrastructure, and participate in on-call rotation for an AI-driven company.
Top Skills: AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIpIscsiJenkinsKubernetesLinux/DebianMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRaidRubyS3ScyllaSshSslSupermicroTcpTlsUbuntu
Reposted One Month AgoSaved
Remote
US
Senior level
Senior level
Big Data • Healthtech • Information Technology • Analytics
Serve as an SRE-focused AI enablement partner: evaluate AI architectures, coach engineering teams on LLM integrations and agentic/RAG patterns, advise on observability, SLOs, governance, and operational readiness, and develop internal standards and documentation.
Top Skills: Agentic FrameworksAnthropic ClaudeAWSAzureAzure Ai FoundryCi/CdDatabricksDatadogDockerGrafanaKubernetesLangchainLlamaindexLlm ApiOpentelemetryRagSemantic Kernel
9 Days AgoSaved
Hybrid
Bellevue, WA, USA
147K-202K Annually
Senior level
147K-202K Annually
Senior level
Cloud
Build and operate reliable, scalable, secure infrastructure for SaaS security and Snowflake data systems. Automate infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and related tools. Ensure security and compliance, participate in on-call rotations, lead incident response and root-cause analysis, and collaborate with development, data science, and security teams on architectural decisions and service implementation.
Top Skills: Artificial IntelligenceCi/CdFlywayInfrastructure As CodeKubernetesMachine LearningSnowflakeSpinnakerTerraform
9 Days AgoSaved
Hybrid
Bellevue, WA, USA
147K-202K Annually
Senior level
147K-202K Annually
Senior level
Cloud
Designs, builds, and operates reliable, scalable infrastructure for security SaaS and Snowflake data systems. Automates infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and Flyway. Partners with security, development, and data science teams on secure, compliant architecture. Participates in on-call rotations, leads critical incident response and root-cause analysis, and implements preventative improvements. In-person onboarding and travel to the Toronto office are required during the first employment week.
Top Skills: Ci/CdContainerizationFlywayInfrastructure As CodeKubernetesSnowflakeSpinnakerTerraform
9 Days AgoSaved
Remote
United States
Senior level
Senior level
Information Technology • Security • Cybersecurity
Own production operations and highly available cloud infrastructure for regulated government environments. Lead incident response, root-cause analysis, disaster recovery testing, compliance operationalization, audit readiness, vulnerability management, and continuous monitoring. Build automation, secure CI/CD pipelines, infrastructure-as-code, observability, and compliance tooling across Kubernetes, Linux, containers, and cloud platforms. Partner with security, compliance, and engineering teams to improve reliability, deployment safety, and regulatory sustainment.
Top Skills: Aws GovcloudBashCi/CdDod Il4Dod Il5FedrampGitopsGoGrafanaKubernetesLinuxNist 800-53PrometheusPythonStigTerraformUnixZero Trust
Reposted 9 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
140K-288K Annually
Senior level
140K-288K Annually
Senior level
Social Media
Operate, scale, and harden an AWS- and Kubernetes-based platform using GitOps. Build CI/CD and infrastructure-as-code (Terraform/Terragrunt), manage ArgoCD/Helm deployments, improve observability, automate toil reduction, lead incident response and post-incident remediation, and partner with application, security, and platform teams to improve reliability and delivery.
Top Skills: ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonRbacTerraformTerragrunt
Reposted 9 Days AgoSaved
In-Office
Raleigh, NC, USA
119K-196K Annually
Senior level
119K-196K Annually
Senior level
Cloud • Information Technology • Internet of Things • Software • Consulting • Infrastructure as a Service (IaaS) • Automation
Design, build, automate, and operate Red Hat Hybrid OpenShift platforms across cloud and on-prem. Implement GitOps, CI/CD, monitoring, and SRE practices; develop operators and tooling; lead incident response, on-call duties, and postmortems; mentor peers and improve platform reliability and self-service.
Top Skills: ArgocdAWSAzureCatchpointCi/CdDatadogDnsFedoraGitlabGitopsGoGoGCPGrafanaHttp/TlsKubernetesKubernetes OperatorsLdapLinuxOpenshiftOpenshift PipelinesOperator SdkPrometheusPythonRhelSplunkSplunk ImTcp/IpTekton
Reposted One Month AgoSaved
Hybrid
Palo Alto, CA, USA
186K-256K Annually
Senior level
186K-256K Annually
Senior level
Automotive • Cloud • Hardware • Software
Design, build, and operate developer platform infrastructure to support firmware build pipelines. Implement IaC patterns, GitOps, Kubernetes operators, CI optimization, cloud reliability, security policy-as-code, and developer tooling while mentoring engineers and collaborating on architecture and risk mitigation.
Top Skills: ArgocdAWSAzureBashCiFluxGCPGitopsGoKubernetesPythonService MeshTerraform
Reposted One Month AgoSaved
Hybrid
2 Locations
30K-120K Annually
Senior level
30K-120K Annually
Senior level
Information Technology • Automation
The SRE/Infrastructure Engineer will architect and manage secure, scalable systems for automated penetration testing, optimizing reliability, and enhancing infrastructure based on customer demand. Responsibilities include maintaining production environments, leading technical discussions, and promoting high coding standards.
Top Skills: AWSAzureCloudFormationElkGCPNew RelicOpentelemetryPostgresPrometheusTerraform
Reposted One Month AgoSaved
Hybrid
Atlanta, GA, USA
Senior level
Senior level
HR Tech • Information Technology • Professional Services • Software • Business Intelligence • Consulting • Automation
Seeking a Site Reliability Engineer with expertise in Unix/Linux, scripting languages, and experience in containerization, cloud platforms, and application monitoring tools.
Top Skills: AnsibleApache TomcatAWSCassandraChefCoradiantDockerDynatraceElasticGomezGCPJenkinsKafkaLinuxMq SeriesOraclePuppetPythonShell ScriptingSplunkTealeafUnixVagrantWebsphere
Reposted One Month AgoSaved
Hybrid
2 Locations
164K-306K Annually
Senior level
164K-306K Annually
Senior level
Software
Operate and improve reliability across Retool Cloud, managed, BYOC, and self-hosted deployments. Automate provisioning, upgrades, migrations, and secret rotations. Build observability and safer deployment/rollback workflows, partner with product teams, and produce runbooks, docs, and migration guides to reduce customer toil and scale operations.
Top Skills: AWSDocker ComposeGoHelmJavaKubernetesPostgresPythonRubyTerraformTypescript
Reposted One Month AgoSaved
In-Office or Remote
8 Locations
170K-230K Annually
Mid level
170K-230K Annually
Mid level
Artificial Intelligence • Cloud • Information Technology • Software
Contribute to the reliability and performance of Mithril's GPU orchestration platform through automation, observability, and infrastructure management. Collaborate with the team to ensure scalability across multi-cloud environments while maintaining systems stability and implementing SLOs.
Top Skills: AWSAzureGCPGoGrafanaKubernetesLinuxOpentelemetryPrometheusPulumiPythonTcp/IpTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
Hybrid
2 Locations
Mid level
Mid level
Fintech • Financial Services
Support and operate OpenShift/Kubernetes platforms and RHEL systems for a trading exchange. Maintain automation (Ansible, Jenkins, ArgoCD, GitOps), cloud workloads (AWS/Azure), authentication/DNS/time services, enterprise hardware/storage, and developer platform integrations. Monitor, respond to incidents, perform patching and changes, and collaborate with developers, traders, and infrastructure teams to ensure security, resilience, and availability.
Top Skills: AnsibleArgocdArtifactoryAWSAzureBashBindChronyCorednsDell PoweredgeGithub ActionsGithub EnterpriseGitopsJenkinsKerberosKubernetesLdapNtpPtpPythonRed Hat Enterprise Linux (Rhel)Red Hat OpenshiftSonarqubeSssdTerraform
Reposted One Month AgoSaved
In-Office
San Francisco, CA, USA
210K-240K Annually
Senior level
210K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
Design, operate, and automate global network and reliability infrastructure for large-scale ML workloads and a private supercomputer. Own device configuration management, protocols (BGP, VPNs, WAN), datacenter fabrics, monitoring/SLOs, incident response, security/compliance, and cross-team reliability improvements.
Top Skills: AirflowAnsibleBashBgpBluefieldCniCumulus LinuxDatadogEcmpElkEvpn/VxlanFirewallsGrafanaInfinibandInfobloxIngressIpsec VpnsIscsiKafkaKubernetesLinuxLoad BalancersLustrefsMplsNetboxNetwork PolicyNfsNornirOpentelemetryPrometheusPythonQosService NetworkingSparkSpectrum-XSpine-LeafSwitchesTerraformVpnsWan Circuits
Reposted One Month AgoSaved
In-Office
St. George, UT, USA
Mid level
Mid level
Cloud
The Site Reliability Engineer at TCN will design, deploy, and maintain systems for performance, reliability, and security, while managing incidents and collaborating with teams.
Top Skills: BashGoGoogle Cloud PlatformJavaKubernetesLinuxNode.jsPythonRuby
Reposted One Month AgoSaved
In-Office
Tyson's Corner, VA, USA
100K-160K Annually
Junior
100K-160K Annually
Junior
Fintech
The Site Reliability Engineer will monitor and manage Kubernetes clusters, optimize Cloud Infrastructure, and automate processes using tools like Terraform and Docker.
Top Skills: Amazon S3AWSAzureC/C++CephDockerGCPHdfsHelmJavaJavaScriptKubernetesNfsPostgresPythonRubyTerraform
Reposted One Month AgoSaved
Remote
USA
Senior level
Senior level
Information Technology • Cryptocurrency
The Site Reliability Engineer will lead technical initiatives, architect solutions, troubleshoot issues, mentor team members, and improve observability practices.
Top Skills: ArgocdBashElk StackGCPGoGrafanaHelmKubernetesPrometheusPythonTerraform
Reposted One Month AgoSaved
Remote
United States
154K-231K Annually
Senior level
154K-231K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Lead technical design and architecture of large-scale Ceph clusters, drive capacity planning and major upgrades, resolve complex cross-functional production incidents, establish automation and reliability standards, and mentor engineers on advanced Ceph operations to scale the global storage platform.
Top Skills: AnsibleBluestoreCephCephfsCinderCrushCsiDdnErasure CodingGoIsilonKubernetesManilaMdsMgrMonNetappNeutronNovaOpenstackOsdPlacement GroupsPurePythonRbdRbd MirroringRgwRgw MultisiteRookS3SaltstackSwiftTerraformVastWeka
Reposted One Month AgoSaved
Remote
United States
128K-192K Annually
Senior level
128K-192K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Own and maintain reliability, performance, scalability, and capacity of massive Ceph storage clusters. Troubleshoot distributed storage issues, build automation (Python/Shell, SaltStack, Ansible), define observability (SLIs/SLOs, PromQL/LogQL, Grafana), and lead lifecycle initiatives like upgrades, expansions, and migrations.
Top Skills: AnsibleCeph (RadosCephfs)ChefGrafanaLinuxLogqlPromqlPuppetPythonRbdRgwSaltstackShell
One Month AgoSaved
In-Office
Los Angeles, CA, USA
100K-200K Annually
Mid level
100K-200K Annually
Mid level
Energy • Chemical • Utilities • Manufacturing
Design, implement, and maintain observability, alerting, and developer productivity systems for production and internal services. Instrument services with metrics, logs, and traces, run on-call, respond to incidents, lead reviews, and automate operational workflows to improve reliability and reduce toil.
Top Skills: Ci/CdDatadogDnsGrafanaHTTPInfrastructure-As-CodeLoggingMetricsOpentelemetryPrometheusTlsTracing
Reposted 10 Days AgoSaved
Hybrid
3 Locations
145K-200K Annually
Senior level
145K-200K Annually
Senior level
Blockchain • Energy • Cryptocurrency
Hands-on role to assess, implement, test, and document backup, restore, failover, and recovery capabilities. Inventory critical systems, design and automate backup and restoration, run recovery exercises, produce runbooks, validate recoverability, measure RTO/RPO, and train system owners. Collaborate with Security, SRE, DevOps, QA, and application teams to harden shared recovery capabilities and transfer operational ownership.
Reposted 10 Days AgoSaved
Hybrid
Bellevue, WA, USA
120K-150K Annually
Senior level
120K-150K Annually
Senior level
Healthtech • Software • Analytics • Business Intelligence
Senior SRE responsible for designing, building, and operating reliable, scalable distributed systems; owning production reliability (SLOs/SLIs, incident response, MTTR reduction); automating toil with software and platform tooling; driving observability, capacity planning, and cross-team reliability improvements; mentoring engineers and running blameless postmortems.
Top Skills: AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Reposted 10 Days AgoSaved
Hybrid
City Point, City of Boston, MA, USA
130K-140K Annually
Senior level
130K-140K Annually
Senior level
Fintech • Software
Senior SRE responsible for ensuring reliability, scalability, and performance of production systems. Investigates and resolves incidents, works with R&D on defects, manages deployments and change validation, builds monitoring and diagnostic tooling to improve MTTA/MTTD/MTTR, configures observability platforms, participates in on-call rotation, and performs risk assessments and production readiness validation.
Top Skills: Ai ToolsAkamaiAmqAnsibleAWSAzureBashCdnCloudflareCloudwatchDatadogDatastreamDnsDockerDynatraceEfsEksHTTPHttpsInterconnectJavaJmsKubernetesMongoDBOraclePostgresPrometheusPythonRabbitMQRedshiftS3SplunkTcpTerraformUdpUnix/LinuxWafZabbix
Reposted 10 Days AgoSaved
In-Office
San Francisco, CA, USA
210K-240K Annually
Senior level
210K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
The Senior Site Reliability Engineer will design and maintain scalable infrastructure, improve system reliability, manage CI/CD pipelines, and collaborate across teams for operational excellence.
Top Skills: AnsibleArgocdAWSBashDatadogDockerElkGithub ActionsGrafanaKubernetesLinuxOpentelemetryPrometheusPythonTerraform
Reposted 10 Days AgoSaved
In-Office
Tyson's Corner, VA, USA
165K-200K Annually
Senior level
165K-200K Annually
Senior level
Fintech
Lead the reliability, monitoring, and scaling of Kubernetes clusters and cloud infrastructure. Deploy and maintain containerized microservices, implement IaC (Terraform/CloudFormation), automate operational tasks, manage incidents, and ensure security and compliance across environments. Collaborate with development and operations teams to optimize performance and availability.
Top Skills: Amazon S3AnsibleApache MesosAWSAzureC/C++CephCloudFormationDockerGCPHdfsHelmJavaJavaScriptJenkinsKubernetesLinuxNfsPostgresPythonRubyTerraformYarn
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account