Top Remote Site Reliability Engineer Jobs

Reposted 16 Days AgoSaved
Remote
United States
Senior level
Senior level
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills: AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Reposted 16 Days AgoSaved
Remote or Hybrid
4 Locations
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills: ArgocdAWSDatadogGitGoKubernetesPythonTerraform
17 Days AgoSaved
Remote
United States
120K-165K Annually
Senior level
120K-165K Annually
Senior level
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills: AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Reposted 18 Days AgoSaved
In-Office or Remote
The Center, IN, USA
180K-250K Annually
Expert/Leader
180K-250K Annually
Expert/Leader
Edtech • Information Technology • Software
Lead infrastructure, reliability, and observability across multi-cloud environments. Improve CI/CD, IaC standards, staging parity, Kubernetes operations, monitoring and SLOs, incident response, and platform modernization while partnering with engineering teams.
Top Skills: Ai-Assisted Development Tools (Claude CodeAutoscalingAWSCi/Cd PipelinesCodex)Event-Driven ArchitecturesGCPIncident ManagementInfrastructure-As-CodeKubernetesMonitoringObservabilityPythonQueue-Based ArchitecturesRuby On RailsSlo FrameworksTerraform
Reposted 18 Days AgoSaved
In-Office or Remote
2 Locations
76K-136K Annually
Mid level
76K-136K Annually
Mid level
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills: Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
Reposted 18 Days AgoSaved
Remote
USA
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills: AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Reposted 18 Days AgoSaved
Remote
United States
180K-224K Annually
Senior level
180K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills: Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Reposted 20 Days AgoSaved
In-Office or Remote
2 Locations
Senior level
Senior level
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills: BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
21 Days AgoSaved
Remote
USA
200K-270K Annually
Expert/Leader
200K-270K Annually
Expert/Leader
Social Media • Software
Design, implement, and operate infrastructure for a federated social network. Own reliability, availability, observability, incident response, deployments, capacity planning, and cost management. Build automation and tooling, scale bare-metal and cloud systems for millions of users, lead incident reviews, mentor engineers, and manage vendor relationships to ensure operational excellence.
Top Skills: Bare-MetalCapacity PlanningCloud ServicesColocationDatabasesDeployment And Rollback SystemsGoIncident ResponseKubernetesLinuxMonitoringNetworkingObservability SystemsProduction AutomationStorage
Reposted 21 Days AgoSaved
Remote or Hybrid
North Carolina, USA
139K-282K Annually
Senior level
139K-282K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Ensure production database services are scalable, resilient, high-performing, and secure. Participate in on-call rotation, monitoring, alerting, incident response, disaster recovery drills, and post-mortems. Automate operational tasks, design reliability/security controls, build observability tooling, capacity forecasts, and collaborate with developers and security to improve hybrid cloud vector and NoSQL database platforms.
Top Skills: AWSAzureCi/CdGCPGitGitJenkinsKubernetesLangchainMilvusMongoDBMySQLNoSQLOpenai ApiOpenshiftPineconePostgresPythonVector DatabasesWeaviate
Reposted 21 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
136K-180K Annually
Senior level
136K-180K Annually
Senior level
Big Data • Energy • Big Data Analytics
The Staff Site Reliability Engineer will lead in designing and maintaining cloud infrastructure on GCP, drive IaC strategy, manage Kubernetes operations, ensure security compliance, and mentor engineers.
Top Skills: BashGoGoogle Cloud PlatformGrafanaKubernetesOpentelemetryPostgresPythonTerraform
Reposted 22 Days AgoSaved
Remote
United States
205K-270K Annually
Senior level
205K-270K Annually
Senior level
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills: ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 22 Days AgoSaved
Remote
United States
Senior level
Senior level
Agency • Information Technology
Lead SRE role designing and maintaining CI/CD pipelines (GitHub Actions), containerized deployments (Docker, Kubernetes, AKS, Helm), web/mobile app releases, observability, automated testing, and DevOps best practices across cloud environments with cross-functional collaboration and regulatory compliance.
Top Skills: AksAndroidAzure Application InsightsAzure Log AnalyticsAzure MonitorBashBranchingDockerDocker ComposeGitGit HooksGithub ActionsGoogle PlayHelmHerokuiOSIos App StoreJavaKubernetesNpmPowershellPull RequestsPythonSonarqubeVeracodeVercel
Reposted 22 Days AgoSaved
Remote
USA
180K-210K Annually
Senior level
180K-210K Annually
Senior level
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills: AWSKubernetesPostgresTerraform
Reposted 22 Days AgoSaved
In-Office or Remote
11 Locations
160K-179K Annually
Senior level
160K-179K Annually
Senior level
Fintech • Payments
The Senior Staff SRE leads reliability engineering initiatives, drives operational excellence, mentors staff, and influences architecture to enhance system reliability and performance.
Top Skills: Ai/MlAWSAzureDockerElk StackGCPGrafanaKubernetesMySQLNoSQLPostgresSplunk
23 Days AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills: AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
23 Days AgoSaved
In-Office or Remote
2 Locations
135K-160K Annually
Senior level
135K-160K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Design, deploy, and maintain AWS-hosted Kubernetes (EKS) infrastructure; build automation and AI-assisted runbooks; provide observability and incident response; ensure HIPAA-compliant, high-availability platform operations and mentor engineering teams.
Top Skills: AWSBashDatadogEc2EksGitGithub ActionsGoHelmKubernetesPythonRdsS3Terraform
23 Days AgoSaved
In-Office or Remote
28 Locations
Senior level
Senior level
Artificial Intelligence • Information Technology • Software
Design, build, and operate Zencoder's cloud and Kubernetes infrastructure to improve reliability, security, scalability, cost efficiency and developer tooling. Automate operations, enhance CI/CD and GitOps, operate data systems (OpenSearch/Postgres), improve observability and incident response, and prepare the platform for growth and AI workloads.
Top Skills: AutoscalingAWSCdnCi/CdDnsElasticsearchFirewallsGCPGitopsGkeIamInfrastructure-As-CodeKubernetesLoad BalancingObservabilityOpensearchPostgresService IdentitySglangVllmVpcWaf
10 Hours AgoSaved
Remote
US
104K-140K Annually
Senior level
104K-140K Annually
Senior level
Software • Financial Services
Own day-to-day AWS and database operations for a serverless production platform, manage backups and disaster recovery, monitor and debug production, lead incident response and on-call, maintain infrastructure-as-code (SST/Pulumi), optimize costs, and mentor the team on operational best practices.
Top Skills: AWSBashCloudwatchDnsDockerEventbridgeGithub ActionsLambdaLinuxPostgresPulumiPythonS3SqsSstTerraformTlsTypescript
Reposted 11 Hours AgoSaved
Remote
USA
125K-165K Annually
Senior level
125K-165K Annually
Senior level
Healthtech
Design, scale, and operate secure AWS cloud infrastructure (EKS, IAM, RBAC); build and maintain IaC (Terraform/Terragrunt), GitHub Actions CI/CD, Datadog observability, and Python automation; document runbooks, participate in on-call rotations, postmortems, and Agile workflows to improve reliability and security.
Top Skills: AWSDatadogEc2EksFargateGithub ActionsGithub Advanced SecurityHelmIamJIRAKubernetesLambdaPythonRbacSecrets ManagerServerlessTerraformTerragruntVpc
Reposted 11 Hours AgoSaved
Remote
United States of America
185K-227K Annually
Senior level
185K-227K Annually
Senior level
Other
The Senior Site Reliability Engineer at Juul Labs ensures operational stability and performance of hybrid cloud infrastructure, leads automation, and handles critical incidents.
Top Skills: AWSBashCloudFormationGCPNutanixPowershellPythonTerraform
Reposted 11 Hours AgoSaved
Remote
US
Senior level
Senior level
Artificial Intelligence
Own operational excellence for cloud infrastructure: run incident management, improve reliability through automation, own a platform domain (e.g., Kubernetes, Temporal, observability), manage vendor and cost relationships, and deliver measurable reductions in incidents and costs within 12 months.
Top Skills: AWSKubernetesLlm ApisMongoDBObservabilityPythonTemporal
11 Hours AgoSaved
In-Office or Remote
2 Locations
168K-270K Annually
Senior level
168K-270K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Drive reliability and observability for GeForce NOW: build SRE tools, support Kubernetes migration with VMI, triage incidents, automate operational tasks, consult on system design and capacity, manage multi-region cloud deployments, and participate in production on‑call rotations to meet service SLOs.
Top Skills: AlertmanagerArgocdAWSAzureBashDatadogGCPGithub ActionsGitlab CiGoKubernetesPrometheusPythonVmi
Reposted YesterdaySaved
Remote
US
160K-200K Annually
Senior level
160K-200K Annually
Senior level
Digital Media • Edtech
Drive reliability and observability of Epic's GCP-based platform. Own cloud infrastructure, container platform (Kubernetes/GKE), CI/CD, observability, and IaC (Terraform). Define SLOs/SLIs, reduce toil, manage security and compliance practices, participate in on-call rotations, lead incident response and post-mortems, and partner with product and data teams to troubleshoot and improve platform reliability.
Top Skills: ArgocdBashCloud MonitoringDockerGceGCPGcsGithub ActionsHelmIamJenkinsKubernetes (Gke)New RelicPythonTerraformVpc
Reposted YesterdaySaved
Remote
United States
Senior level
Senior level
Logistics • Software
Own and operate scalable infrastructure on GCP (GKE, Cloud Run, AlloyDB); author Terraform modules; manage containerized workloads; build observability in Datadog; design CI/CD in GitHub Actions; automate operational workflows; lead incident response and post-mortems; partner with engineers to improve reliability, cost, and automation.
Top Skills: AlloydbClickhouseCloud RunCloudflare WorkersDatadogDockerGCPGithub ActionsGkeGoGrafanaIamKafkaKubernetesNetworkingPostgresPrometheusPub/SubPythonRedisRedpandaTerraformTypescript
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account