Top Remote Site Reliability Engineer Jobs

11 Days AgoSaved
In-Office or Remote
2 Locations
147K-264K Annually
Senior level
147K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale distributed systems.
Top Skills: Cloud ComputingDistributed SystemsHTTPJavaLinux/UnixPerlPythonTcp/IpTls/Ssl
Reposted 13 Days AgoSaved
In-Office or Remote
8 Locations
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills: AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Reposted 13 Days AgoSaved
In-Office or Remote
2 Locations
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills: AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Reposted 13 Days AgoSaved
Remote or Hybrid
Location, WV, USA
Senior level
Senior level
Payments
Senior SRE responsible for ensuring high availability and resiliency of a global payments platform by building observability, automations, AI-driven remediation, incident response, and self-healing workflows; participates in on-call rotation and hybrid Philadelphia-based work.
Top Skills: AiopsAksAnthropic (Claude)ApmAzureAzure Ai (Foundry)Azure Sre AgentCi/CdDatadogDnsDynatraceHTTPHttpsIisKubernetesLoad BalancingNew RelicOpenai (Codex)Pagerduty Process AutomationPowershellPythonRundeckSQLT-SqlTcp/IpVMwareWindows Server
Reposted 15 Days AgoSaved
Remote
2 Locations
150K-195K Annually
Senior level
150K-195K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills: DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Reposted 15 Days AgoSaved
In-Office or Remote
Washington, DC, USA
113K-162K Annually
Senior level
113K-162K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Lead reliability and implementation for Kong's Managed Gateways: design and operate multi-cloud, Kubernetes-based systems, own incident response and SLOs, automate CI/CD and IaC, mentor SREs, and drive enterprise customer onboarding and technical implementations.
Top Skills: AnsibleAWSAzureCassandraCi/CdDatadogElkGCPGoGrafanaIstioKubernetesLinkerdPostgresPrometheusTerraform
Reposted 15 Days AgoSaved
Remote
USA
117K-181K Annually
Senior level
117K-181K Annually
Senior level
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills: AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Reposted 15 Days AgoSaved
In-Office or Remote
7 Locations
Senior level
Senior level
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills: AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Reposted 16 Days AgoSaved
Remote
USA
160K-208K Annually
Senior level
160K-208K Annually
Senior level
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills: AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Reposted 17 Days AgoSaved
Remote or Hybrid
Redmond, WA, USA
120K-150K Annually
Senior level
120K-150K Annually
Senior level
Healthtech • Software • Analytics • Business Intelligence
Lead and own reliability for critical backend and distributed systems: design, launch, on-call, incident leadership, SLO/SLI/error budget definition, automation to remove toil, observability improvement, resilience testing, mentoring, and cross-team reliability initiatives for production healthcare workflows.
Top Skills: AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Reposted 17 Days AgoSaved
Remote
US
Senior level
Senior level
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills: BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Reposted 17 Days AgoSaved
In-Office or Remote
Norfolk, VA, USA
Senior level
Senior level
Information Technology
The Senior Site Reliability Engineer is responsible for architecting reliability strategies, implementing SRE frameworks, mentoring engineers, and ensuring system resilience and performance in government systems.
Top Skills: Cloud ArchitectureDevsecopsGoInfrastructure As Code (Iac)JavaKubernetesLinuxNist 800-53PythonRmf
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 17 Days AgoSaved
Remote
United States
Senior level
Senior level
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills: ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
Reposted 17 Days AgoSaved
Remote
United States
160K-180K Annually
Senior level
160K-180K Annually
Senior level
Software
Own and improve platform performance, reliability, and deployment automation. Manage cloud infrastructure, implement IaC, monitor systems with observability tools, provide operational support for distributed applications, and integrate production learnings into development workflows.
Top Skills: Aiops ToolingAws Elastic ContainersAws RdsAws S3Claude CodeClaude CoworkDatadogHarness EngineeringInfrastructure As CodeKubernetesLlmsPrompt EngineeringRigorSplunk
18 Days AgoSaved
Remote or Hybrid
3 Locations
240K-356K Annually
Senior level
240K-356K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Operate and scale bare-metal Kubernetes clusters at large scale, build control-plane services, operators, and automation for cluster lifecycle. Automate tooling in Go and Python, define SLOs/SLIs, run on-call rotations, troubleshoot incidents, and support customers integrating workloads, storage, and authentication.
Top Skills: ArgocdCi/CdCluster ApiCniCrdCsiEksFluentbitGitopsGkeGoGrafanaHelmKubeadmKubernetesKubernetes OperatorsPrometheusPython
18 Days AgoSaved
Remote or Hybrid
3 Locations
240K-356K Annually
Senior level
240K-356K Annually
Senior level
Software
Operate and scale bare-metal Kubernetes clusters to thousands of nodes, manage incidents and on-call rotation, support customers, collaborate with HPC/datacenter teams, build automation and tooling in Go/Python, create control-plane services/operators, automate cluster lifecycle, and define SLOs/SLIs for platform reliability.
Top Skills: ArgocdBare-Metal KubernetesCi/CdCluster ApiCniCrdCsiEksFluentbitGkeGoGpuGrafanaHelmKubeadmKubernetesKubernetes OperatorsLinuxPrometheusPython
Reposted 18 Days AgoSaved
Remote
United States
175K-195K Annually
Senior level
175K-195K Annually
Senior level
Legal Tech • Software
Design and improve observability (monitoring, logging, tracing, SLIs/SLOs), build automation and CI/CD, lead incident response and reliability improvements, mentor SREs, run on-call, and apply AI/ML to operational signals to forecast and reduce risks.
Top Skills: AWSBashCi/CdDistributed TracingGoInfrastructure As CodeKubernetesLoggingMonitoringPythonSlisSlos
Reposted 19 Days AgoSaved
In-Office or Remote
Plano, TX, USA
117K-209K Annually
Senior level
117K-209K Annually
Senior level
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for production services in Autodesk GovCloud by building automation, SLO/SLI practices, observability, incident response, resilience testing, and runbooks. Deploy, operate, and improve cloud services while ensuring compliance (FedRAMP) and participating in 24x7 on-call rotations and cross-team collaboration.
Top Skills: APIsAWSAws GovcloudAzureBashCaching TechnologiesCi/CdCloudwatchContainersDatabasesDatadogDeployment AutomationDistributed SystemsDnsDynatraceGoInfrastructure As CodeJavaKubernetesLoad BalancingMessaging SystemsNetworkingPowershellPythonSplunkStorage Platforms
Reposted 19 Days AgoSaved
Remote
Idaho, USA
117K-209K Annually
Senior level
117K-209K Annually
Senior level
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for production services in Autodesk GovCloud: deploy, operate, and automate cloud services; define SLOs/SLIs and observability; drive incident response, resilience testing, and toil reduction; ensure compliance (FedRAMP) and participate in 24x7 on-call rotation.
Top Skills: APIsAWSAws GovcloudAzureBashCi/CdCloudwatchContainersDatadogDnsDynatraceGoInfrastructure As CodeJavaKubernetesLoad BalancingNetworkingPowershellPythonSplunk
Reposted 19 Days AgoSaved
Remote
United States
141K-208K Annually
Senior level
141K-208K Annually
Senior level
Database • Analytics
This role involves ensuring the reliability and performance of ClickHouse's cloud infrastructure, collaborating with engineering teams, incident management, and driving continuous improvement in service availability.
Top Skills: AnsibleAWSAzureClickhouseDocker SwarmGoGoogle Cloud PlatformKubernetesPuppetPythonTerraform
Reposted 21 Days AgoSaved
Remote
United States
165K-195K Annually
Senior level
165K-195K Annually
Senior level
Fintech • Real Estate • Software
Lead reliability and observability efforts across the org: design and maintain Kubernetes and AWS infrastructure, build CI/CD pipelines, drive IaC standards (Terraform/Crossplane), partner with 16+ teams to roll out tools and processes, participate in on-call rotation and incident response, and use AI tools to accelerate work.
Top Skills: Ai ToolsArgoAurora PostgresAWSCi/Cd PipelinesCrossplaneDatadogDocumentdb (Mongo)EcsEksGithub ActionsHelmKubernetesMongoDBPostgresRdsTerraform
21 Days AgoSaved
In-Office or Remote
4 Locations
105K-192K Annually
Senior level
105K-192K Annually
Senior level
Information Technology • Legal Tech • Analytics
Partner with DBAs and SRE teams to ensure reliable, secure, and scalable database infrastructure. Lead incident management, vulnerability remediation, DR planning and resilience testing, infrastructure provisioning across VMs and Chainguard containers, cost optimization, and SRE capability building. Provide on-call support and coach junior staff.
Top Skills: Amazon Web Services (Aws)ChainguardGitlabGrafanaIaasJenkinsLinuxMySQLPmm3PostgresPrometheusTerraformVm
21 Days AgoSaved
Remote
US
104K-163K Annually
Senior level
104K-163K Annually
Senior level
Software • Financial Services
Own day-to-day AWS and database operations for a serverless production platform, manage backups and disaster recovery, monitor and debug production, lead incident response and on-call, maintain infrastructure-as-code (SST/Pulumi), optimize costs, and mentor the team on operational best practices.
Top Skills: AWSBashCloudwatchDnsDockerEventbridgeGithub ActionsLambdaLinuxPostgresPulumiPythonS3SqsSstTerraformTlsTypescript
Reposted 21 Days AgoSaved
Remote
USA
125K-165K Annually
Senior level
125K-165K Annually
Senior level
Healthtech
Design, scale, and operate secure AWS cloud infrastructure (EKS, IAM, RBAC); build and maintain IaC (Terraform/Terragrunt), GitHub Actions CI/CD, Datadog observability, and Python automation; document runbooks, participate in on-call rotations, postmortems, and Agile workflows to improve reliability and security.
Top Skills: AWSDatadogEc2EksFargateGithub ActionsGithub Advanced SecurityHelmIamJIRAKubernetesLambdaPythonRbacSecrets ManagerServerlessTerraformTerragruntVpc
Reposted 21 Days AgoSaved
Remote
United States of America
185K-227K Annually
Senior level
185K-227K Annually
Senior level
Other
The Senior Site Reliability Engineer at Juul Labs ensures operational stability and performance of hybrid cloud infrastructure, leads automation, and handles critical incidents.
Top Skills: AWSBashCloudFormationGCPNutanixPowershellPythonTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account