Top Site Reliability Engineer Jobs

6 Days AgoSaved
In-Office or Remote
2 Locations
Senior level
Senior level
Energy
Build, transition, operate, and improve a multi-cloud production platform across Azure, GCP, AWS, and Kubernetes. Responsibilities include infrastructure as code, GitOps and CI/CD, managed Kubernetes operations, troubleshooting, incident response, observability, reliability improvements, disaster recovery, automation, and cost control. The role partners with architects and global teams to ensure platforms are secure, scalable, recoverable, and supportable.
Top Skills: AksAnsibleAWSAzureCi/CdEksGitopsGkeGoGoogle Cloud PlatformKubernetesLinuxMongoDBMySQLNetworkingOpentofuPostgresPythonShell ScriptingTerraformTerragrunt
Reposted 28 Days AgoSaved
In-Office
San Francisco, CA, USA
200K-275K Annually
Senior level
200K-275K Annually
Senior level
Artificial Intelligence • Healthtech • Information Technology • Software
As a Site Reliability Engineer, you will manage the production environment, focusing on infrastructure design, automation, and optimizing deployment pipelines to ensure high availability.
Top Skills: HelmKafkaKubernetesPostgresPythonRedisTerraformTypescript
Reposted 28 Days AgoSaved
In-Office
Bedford, NH, USA
Senior level
Senior level
Agency • Marketing Tech • Software • Consulting
Lead and maintain performance, security, and reliability of client hosting environments across multi-cloud platforms. Architect resilient infrastructure, manage IaC and CI/CD, administer Windows/IIS and WP Engine environments, handle SSL/DNS/SSO, participate in on-call rotations, and engage with clients as senior escalation and trusted advisor.
Top Skills: App ServicesApplication InsightsAWSAzureAzure DevopsAzure SqlCi/CdDnsEc2IamIisKey VaultsPowershellRdsS3Ssl/TlsSsoTerraformVulnerability ManagementWindows ServerWp Engine
6 Days AgoSaved
In-Office
Culver City, CA, USA
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Digital Media • Software
Deploy, operate, validate, and improve Ateme video delivery platforms in customer environments, including critical live production systems. Responsibilities include Linux administration, networking, video technology integration, troubleshooting, performance investigations, documentation, customer training, technical collaboration with R&D, and ensuring service availability. The role supports live production events and participates in a 24/7 on-call rotation.
Top Skills: ArqAutomationAv1AvcCdnCloud InfrastructureCmafDistributed SystemsDrmHevcHlsJpeg-XsKubernetesLinuxMpeg-TsNetworkingObservabilityScte-104Scte-224Scte-30Scte-35Smpte 2110UnixVideo Encoding
6 Days AgoSaved
In-Office
2 Locations
128K-216K Annually
Senior level
128K-216K Annually
Senior level
eCommerce • Fintech • Information Technology • Payments • Financial Services
Manage, deploy, and architect highly scalable, redundant infrastructure across AWS and Kubernetes environments. Build CI/CD pipelines with GitHub Actions, automate infrastructure using Terraform, manage databases and document storage, and implement observability with monitoring and tracing tools. Lead technical design reviews, mentor teams on SRE principles, and improve software performance through infrastructure architecture.
Top Skills: AWSBashDatadogDnsDockerDocument StorageDynatraceGitGithub ActionsKubernetesLinuxLoad BalancingNew RelicNode.jsPythonRdbmsRuby On RailsTerraformUnixVirtual Networking
6 Days AgoSaved
Remote or Hybrid
United States
155K-172K Annually
Senior level
155K-172K Annually
Senior level
Healthtech • Software
Provides technical leadership for site reliability and operational development. Designs automated, repeatable cloud and infrastructure solutions; monitors service objectives and business metrics; improves application resiliency, performance, efficiency, and cost. Builds monitoring, deployment, testing, and vulnerability-response automation while supporting mission-critical production systems. Collaborates with software, security, product, and business teams, conducts technical training and resilience exercises, and promotes sound development, change-management, and operational practices.
Top Skills: Amazon EcsAnsibleAzureCC++ChefDockerGoJavaKubernetesLinuxPerlPuppetPythonRubyTerraformWindows
Reposted 11 Days AgoSaved
In-Office
Seattle, WA, USA
160K-250K Annually
Mid level
160K-250K Annually
Mid level
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, optimize workflows for teams, manage secure infrastructure, and participate in on-call duties.
Top Skills: AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIscsiJenkinsKubernetesLinux/Debian Family/UbuntuMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRubyS3ScyllaSshSslSupermicroTcpTls
29 Days AgoSaved
In-Office
2 Locations
106K-145K Annually
Mid level
106K-145K Annually
Mid level
Hardware • Information Technology • Other • Software • Analytics
Serve as primary technical escalation for enterprise cloud customers, design and automate deployment pipelines, ensure Azure integrations, implement Terraform IaC, manage CI/CD with GitHub Actions/Azure DevOps, and use New Relic and AI-driven monitoring to improve reliability and performance.
Top Skills: Azure DevopsBashCi/CdGithub ActionsAzureNew RelicPowershellTerraform
Reposted 29 Days AgoSaved
In-Office
Lovelace, NC, USA
Senior level
Senior level
Artificial Intelligence • Machine Learning • Security • Database • Analytics • Big Data Analytics
As a Site Reliability Engineer, you'll ensure the availability and performance of AI applications, maintain infrastructure, automate tasks, and troubleshoot issues in high-scale environments.
Top Skills: AnsibleAWSAzureBashCircleCICloudFormationDatadogDockerDynatraceEc2Elk StackGCPGitlab CiGoGrafanaJenkinsKubernetesLambdaLinuxPrometheusPythonS3TerraformUnix
Reposted 29 Days AgoSaved
Hybrid
Santa Clara, CA, USA
175K-215K Annually
Senior level
175K-215K Annually
Senior level
Artificial Intelligence • Robotics • Automation • Manufacturing
Lead the Platform & SRE team to design and operate a unified deployment platform spanning cloud and on-premise. Architect Pulumi/Kubernetes-based deployments, package software for industrial hardware, build GitHub Actions CI/CD (including HIL), and define observability with Prometheus, Grafana, and OIDC.
Top Skills: AWSAzureCertificate ManagementCnisGCPGithub ActionsGoGrafanaHelmHilKubernetesLinux NetworkingMulti-Cluster ManagementOidcPrometheusPulumiTerraform
Reposted 29 Days AgoSaved
In-Office
2 Locations
192K-455K Annually
Senior level
192K-455K Annually
Senior level
Artificial Intelligence • Information Technology • Cybersecurity • Defense
The Forward Deployed Site Reliability Engineer ensures the reliability of a mission-critical platform, manages incident response, defines SLIs and SLOs, and liaises between engineering and government customers.
Top Skills: AWSBashDockerGrafanaLokiMimirPrometheusPythonTerraform
Reposted 29 Days AgoSaved
Hybrid
Emeryville, CA, USA
Entry level
Entry level
Artificial Intelligence • Machine Learning • Biotech • Generative AI
The Site Reliability Engineer will manage digital infrastructure, ensuring access to compute resources, automating processes, and maintaining resource visibility for researchers.
Top Skills: AnsibleDockerGrafanaKubernetesPrometheusPythonTailscaleTalos Linux
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 29 Days AgoSaved
In-Office
San Francisco, CA, USA
200K-260K Annually
Expert/Leader
200K-260K Annually
Expert/Leader
Big Data
Lead and operate cloud infrastructure and SRE practices, focusing on IaC, CI/CD, observability, incident response, and production-grade agentic AI/LLM systems. Architect, build, and run scalable platforms, drive reliability and SLOs, author runbooks/ADRs, mentor engineers, participate in on-call rotations, and reduce toil via automation and platform tooling.
Top Skills: Agent OrchestrationAgentic AiAWSAzureDockerFluxcdGCPGithub ActionsGitopsGoGrafanaJavaJenkinsKubernetesLinuxLlmLlm GatewayLokiOpentofuPrometheusPythonSentrySignozTcp/IpTerraform
Reposted 29 Days AgoSaved
In-Office or Remote
3 Locations
Senior level
Senior level
Artificial Intelligence
The Deployment Engineer will build and operate AI inference clusters, ensure scalable deployments, optimize allocation, and maintain infrastructure. Responsibilities include software updates, telemetry development, and collaborative improvements with teams.
Top Skills: DockerGrafanaInfluxdbK8SLinuxPrometheusPython
7 Days AgoSaved
Hybrid
New York, NY, USA
Senior level
Senior level
AdTech • Big Data • Internet of Things • Marketing Tech • Mobile • Software • Analytics
Senior Site Reliability Engineer responsible for improving platform availability, scalability, resilience, and observability. The role manages AWS and EKS infrastructure with Terraform, builds CI/CD pipelines using GitHub Actions, expands GitOps through ArgoCD and Helm, strengthens disaster recovery, supports incident response, and develops monitoring, alerting, dashboards, and SLOs. The engineer partners with product teams, leads infrastructure initiatives, troubleshoots Kubernetes, and supports high-availability distributed systems.
Top Skills: Amazon EksAmazon RdsAmazon VpcArgocdAWSAws IamCluster AutoscalerCrossplaneDatadogGithub ActionsGoHelmKarpenterKubernetesOpentofuPrometheusPythonShell ScriptingTerraform
One Month AgoSaved
Remote
United States
152K-253K Annually
Senior level
152K-253K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Build and operate the Veeam Data Cloud GOV environment: map systems, write runbooks, define SLIs/SLOs, run incident response, close observability gaps, automate deployments and support fleet management while working across security and compliance constraints.
Top Skills: Api ManagementApplication InsightsArgocdAws CloudformationAzureAzure Arm TemplatesAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
Reposted 11 Days AgoSaved
In-Office
San Francisco, CA, USA
160K-250K Annually
Mid level
160K-250K Annually
Mid level
Artificial Intelligence • Cloud • Software
The Senior Site Reliability Engineer will automate operations, improve workflows, manage secure infrastructure, and participate in on-call rotation for an AI-driven company.
Top Skills: AristaAWSBashCephChefCifsCiscoDnsDockerElk StackFortinetHpHTTPIcmpIpIscsiJenkinsKubernetesLinux/DebianMesosphereNfsNode.jsPivotal GreenplumPostgresPythonRabbitMQRaidRubyS3ScyllaSshSslSupermicroTcpTlsUbuntu
Reposted One Month AgoSaved
Remote
US
Senior level
Senior level
Big Data • Healthtech • Information Technology • Analytics
Serve as an SRE-focused AI enablement partner: evaluate AI architectures, coach engineering teams on LLM integrations and agentic/RAG patterns, advise on observability, SLOs, governance, and operational readiness, and develop internal standards and documentation.
Top Skills: Agentic FrameworksAnthropic ClaudeAWSAzureAzure Ai FoundryCi/CdDatabricksDatadogDockerGrafanaKubernetesLangchainLlamaindexLlm ApiOpentelemetryRagSemantic Kernel
8 Days AgoSaved
Hybrid
Bellevue, WA, USA
147K-202K Annually
Senior level
147K-202K Annually
Senior level
Cloud
Build and operate reliable, scalable, secure infrastructure for SaaS security and Snowflake data systems. Automate infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and related tools. Ensure security and compliance, participate in on-call rotations, lead incident response and root-cause analysis, and collaborate with development, data science, and security teams on architectural decisions and service implementation.
Top Skills: Artificial IntelligenceCi/CdFlywayInfrastructure As CodeKubernetesMachine LearningSnowflakeSpinnakerTerraform
8 Days AgoSaved
Hybrid
Bellevue, WA, USA
147K-202K Annually
Senior level
147K-202K Annually
Senior level
Cloud
Designs, builds, and operates reliable, scalable infrastructure for security SaaS and Snowflake data systems. Automates infrastructure provisioning, deployments, and incident response using Terraform, Spinnaker, Kubernetes, and Flyway. Partners with security, development, and data science teams on secure, compliant architecture. Participates in on-call rotations, leads critical incident response and root-cause analysis, and implements preventative improvements. In-person onboarding and travel to the Toronto office are required during the first employment week.
Top Skills: Ci/CdContainerizationFlywayInfrastructure As CodeKubernetesSnowflakeSpinnakerTerraform
8 Days AgoSaved
Remote
United States
Senior level
Senior level
Information Technology • Security • Cybersecurity
Own production operations and highly available cloud infrastructure for regulated government environments. Lead incident response, root-cause analysis, disaster recovery testing, compliance operationalization, audit readiness, vulnerability management, and continuous monitoring. Build automation, secure CI/CD pipelines, infrastructure-as-code, observability, and compliance tooling across Kubernetes, Linux, containers, and cloud platforms. Partner with security, compliance, and engineering teams to improve reliability, deployment safety, and regulatory sustainment.
Top Skills: Aws GovcloudBashCi/CdDod Il4Dod Il5FedrampGitopsGoGrafanaKubernetesLinuxNist 800-53PrometheusPythonStigTerraformUnixZero Trust
Reposted 8 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
140K-288K Annually
Senior level
140K-288K Annually
Senior level
Social Media
Operate, scale, and harden an AWS- and Kubernetes-based platform using GitOps. Build CI/CD and infrastructure-as-code (Terraform/Terragrunt), manage ArgoCD/Helm deployments, improve observability, automate toil reduction, lead incident response and post-incident remediation, and partner with application, security, and platform teams to improve reliability and delivery.
Top Skills: ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonRbacTerraformTerragrunt
Reposted 8 Days AgoSaved
In-Office
Raleigh, NC, USA
119K-196K Annually
Senior level
119K-196K Annually
Senior level
Cloud • Information Technology • Internet of Things • Software • Consulting • Infrastructure as a Service (IaaS) • Automation
Design, build, automate, and operate Red Hat Hybrid OpenShift platforms across cloud and on-prem. Implement GitOps, CI/CD, monitoring, and SRE practices; develop operators and tooling; lead incident response, on-call duties, and postmortems; mentor peers and improve platform reliability and self-service.
Top Skills: ArgocdAWSAzureCatchpointCi/CdDatadogDnsFedoraGitlabGitopsGoGoGCPGrafanaHttp/TlsKubernetesKubernetes OperatorsLdapLinuxOpenshiftOpenshift PipelinesOperator SdkPrometheusPythonRhelSplunkSplunk ImTcp/IpTekton
Reposted One Month AgoSaved
Hybrid
Palo Alto, CA, USA
186K-256K Annually
Senior level
186K-256K Annually
Senior level
Automotive • Cloud • Hardware • Software
Design, build, and operate developer platform infrastructure to support firmware build pipelines. Implement IaC patterns, GitOps, Kubernetes operators, CI optimization, cloud reliability, security policy-as-code, and developer tooling while mentoring engineers and collaborating on architecture and risk mitigation.
Top Skills: ArgocdAWSAzureBashCiFluxGCPGitopsGoKubernetesPythonService MeshTerraform
Reposted One Month AgoSaved
Hybrid
2 Locations
30K-120K Annually
Senior level
30K-120K Annually
Senior level
Information Technology • Automation
The SRE/Infrastructure Engineer will architect and manage secure, scalable systems for automated penetration testing, optimizing reliability, and enhancing infrastructure based on customer demand. Responsibilities include maintaining production environments, leading technical discussions, and promoting high coding standards.
Top Skills: AWSAzureCloudFormationElkGCPNew RelicOpentelemetryPostgresPrometheusTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account