Maximum of 25 job preferences reached.
Top Remote Site Reliability Engineer Jobs
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale distributed systems.
Top Skills:
Cloud ComputingDistributed SystemsHTTPJavaLinux/UnixPerlPythonTcp/IpTls/Ssl
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills:
AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Payments
Senior SRE responsible for ensuring high availability and resiliency of a global payments platform by building observability, automations, AI-driven remediation, incident response, and self-healing workflows; participates in on-call rotation and hybrid Philadelphia-based work.
Top Skills:
AiopsAksAnthropic (Claude)ApmAzureAzure Ai (Foundry)Azure Sre AgentCi/CdDatadogDnsDynatraceHTTPHttpsIisKubernetesLoad BalancingNew RelicOpenai (Codex)Pagerduty Process AutomationPowershellPythonRundeckSQLT-SqlTcp/IpVMwareWindows Server
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills:
DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Lead reliability and implementation for Kong's Managed Gateways: design and operate multi-cloud, Kubernetes-based systems, own incident response and SLOs, automate CI/CD and IaC, mentor SREs, and drive enterprise customer onboarding and technical implementations.
Top Skills:
AnsibleAWSAzureCassandraCi/CdDatadogElkGCPGoGrafanaIstioKubernetesLinkerdPostgresPrometheusTerraform
Reposted 15 Days AgoSaved
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills:
AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills:
AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills:
AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Healthtech • Software • Analytics • Business Intelligence
Lead and own reliability for critical backend and distributed systems: design, launch, on-call, incident leadership, SLO/SLI/error budget definition, automation to remove toil, observability improvement, resilience testing, mentoring, and cross-team reliability initiatives for production healthcare workflows.
Top Skills:
AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills:
BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Information Technology
The Senior Site Reliability Engineer is responsible for architecting reliability strategies, implementing SRE frameworks, mentoring engineers, and ensuring system resilience and performance in government systems.
Top Skills:
Cloud ArchitectureDevsecopsGoInfrastructure As Code (Iac)JavaKubernetesLinuxNist 800-53PythonRmf
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills:
ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
Software
Own and improve platform performance, reliability, and deployment automation. Manage cloud infrastructure, implement IaC, monitor systems with observability tools, provide operational support for distributed applications, and integrate production learnings into development workflows.
Top Skills:
Aiops ToolingAws Elastic ContainersAws RdsAws S3Claude CodeClaude CoworkDatadogHarness EngineeringInfrastructure As CodeKubernetesLlmsPrompt EngineeringRigorSplunk
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Operate and scale bare-metal Kubernetes clusters at large scale, build control-plane services, operators, and automation for cluster lifecycle. Automate tooling in Go and Python, define SLOs/SLIs, run on-call rotations, troubleshoot incidents, and support customers integrating workloads, storage, and authentication.
Top Skills:
ArgocdCi/CdCluster ApiCniCrdCsiEksFluentbitGitopsGkeGoGrafanaHelmKubeadmKubernetesKubernetes OperatorsPrometheusPython
Software
Operate and scale bare-metal Kubernetes clusters to thousands of nodes, manage incidents and on-call rotation, support customers, collaborate with HPC/datacenter teams, build automation and tooling in Go/Python, create control-plane services/operators, automate cluster lifecycle, and define SLOs/SLIs for platform reliability.
Top Skills:
ArgocdBare-Metal KubernetesCi/CdCluster ApiCniCrdCsiEksFluentbitGkeGoGpuGrafanaHelmKubeadmKubernetesKubernetes OperatorsLinuxPrometheusPython
Legal Tech • Software
Design and improve observability (monitoring, logging, tracing, SLIs/SLOs), build automation and CI/CD, lead incident response and reliability improvements, mentor SREs, run on-call, and apply AI/ML to operational signals to forecast and reduce risks.
Top Skills:
AWSBashCi/CdDistributed TracingGoInfrastructure As CodeKubernetesLoggingMonitoringPythonSlisSlos
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for production services in Autodesk GovCloud by building automation, SLO/SLI practices, observability, incident response, resilience testing, and runbooks. Deploy, operate, and improve cloud services while ensuring compliance (FedRAMP) and participating in 24x7 on-call rotations and cross-team collaboration.
Top Skills:
APIsAWSAws GovcloudAzureBashCaching TechnologiesCi/CdCloudwatchContainersDatabasesDatadogDeployment AutomationDistributed SystemsDnsDynatraceGoInfrastructure As CodeJavaKubernetesLoad BalancingMessaging SystemsNetworkingPowershellPythonSplunkStorage Platforms
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Lead reliability for production services in Autodesk GovCloud: deploy, operate, and automate cloud services; define SLOs/SLIs and observability; drive incident response, resilience testing, and toil reduction; ensure compliance (FedRAMP) and participate in 24x7 on-call rotation.
Top Skills:
APIsAWSAws GovcloudAzureBashCi/CdCloudwatchContainersDatadogDnsDynatraceGoInfrastructure As CodeJavaKubernetesLoad BalancingNetworkingPowershellPythonSplunk
Database • Analytics
This role involves ensuring the reliability and performance of ClickHouse's cloud infrastructure, collaborating with engineering teams, incident management, and driving continuous improvement in service availability.
Top Skills:
AnsibleAWSAzureClickhouseDocker SwarmGoGoogle Cloud PlatformKubernetesPuppetPythonTerraform
Fintech • Real Estate • Software
Lead reliability and observability efforts across the org: design and maintain Kubernetes and AWS infrastructure, build CI/CD pipelines, drive IaC standards (Terraform/Crossplane), partner with 16+ teams to roll out tools and processes, participate in on-call rotation and incident response, and use AI tools to accelerate work.
Top Skills:
Ai ToolsArgoAurora PostgresAWSCi/Cd PipelinesCrossplaneDatadogDocumentdb (Mongo)EcsEksGithub ActionsHelmKubernetesMongoDBPostgresRdsTerraform
Information Technology • Legal Tech • Analytics
Partner with DBAs and SRE teams to ensure reliable, secure, and scalable database infrastructure. Lead incident management, vulnerability remediation, DR planning and resilience testing, infrastructure provisioning across VMs and Chainguard containers, cost optimization, and SRE capability building. Provide on-call support and coach junior staff.
Top Skills:
Amazon Web Services (Aws)ChainguardGitlabGrafanaIaasJenkinsLinuxMySQLPmm3PostgresPrometheusTerraformVm
Software • Financial Services
Own day-to-day AWS and database operations for a serverless production platform, manage backups and disaster recovery, monitor and debug production, lead incident response and on-call, maintain infrastructure-as-code (SST/Pulumi), optimize costs, and mentor the team on operational best practices.
Top Skills:
AWSBashCloudwatchDnsDockerEventbridgeGithub ActionsLambdaLinuxPostgresPulumiPythonS3SqsSstTerraformTlsTypescript
Healthtech
Design, scale, and operate secure AWS cloud infrastructure (EKS, IAM, RBAC); build and maintain IaC (Terraform/Terragrunt), GitHub Actions CI/CD, Datadog observability, and Python automation; document runbooks, participate in on-call rotations, postmortems, and Agile workflows to improve reliability and security.
Top Skills:
AWSDatadogEc2EksFargateGithub ActionsGithub Advanced SecurityHelmIamJIRAKubernetesLambdaPythonRbacSecrets ManagerServerlessTerraformTerragruntVpc
Other
The Senior Site Reliability Engineer at Juul Labs ensures operational stability and performance of hybrid cloud infrastructure, leads automation, and handles critical incidents.
Top Skills:
AWSBashCloudFormationGCPNutanixPowershellPythonTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Remote Site Reliability Engineers
See AllPopular Job Searches
All Remote Software Engineer Jobs
Remote .NET Developer Jobs
Remote AI Engineer Jobs
Remote Android Developer Jobs
Remote Android Engineer Jobs
Remote Automation Engineer Jobs
Remote AWS Jobs
Remote Backend Engineer Jobs
Remote C# Jobs
Remote C++ Jobs
Remote Cloud Architect Jobs
Remote Cloud Engineer Jobs
Remote Design Engineer Jobs
Remote DevOps Engineer Jobs
Remote DevOps Jobs
Remote Embedded Software Engineer Jobs
Remote Engineering Director Jobs
Remote Engineering Manager Jobs
Remote Enterprise Architect Jobs
Remote Field Engineer Jobs
Remote Front-End Developer Jobs
Remote Front-End Engineer Jobs
Remote Full-Stack Engineer Jobs
Remote Game Developer Jobs
Remote Golang Jobs
Remote Hardware Engineer Jobs
Remote Infrastructure Engineer Jobs
Remote Integration Engineer Jobs
Remote iOS Developer Jobs
Remote iOS Engineer Jobs
Remote IT Engineer Jobs
Remote Java Developer Jobs
Remote Javascript Jobs
Remote Lead Software Engineer Jobs
Remote Linux Engineer Jobs
Remote Linux Jobs
Remote Network Engineer Jobs
Remote Perl Jobs
Remote PHP Developer Jobs
Remote Platform Engineer Jobs
Remote Principal Software Engineer Jobs
Remote Project Engineer Jobs
Remote Python Developer + Engineer Jobs
Remote Python Jobs
Remote QA Analyst Jobs
Remote QA Automation Engineer Jobs
Remote QA Engineer Jobs
Remote Ruby Jobs
Remote Sales Engineer Jobs
Remote Salesforce Administrator Jobs
Remote Salesforce Developer Jobs
Remote Salesforce Developer Jobs
Remote Scala Jobs
Remote Senior DevOps Engineer Jobs
Remote Software Architect Jobs
Remote Software Development Manager Jobs
Remote Software Engineering Manager Jobs
Remote Solutions Architect Jobs
Remote Solutions Engineer Jobs
Remote SRE Jobs
Remote Staff Software Engineer Jobs
Remote Systems Engineer Jobs
Remote Tech Lead Jobs
Remote Test Engineer Jobs
Remote VP of Engineering Jobs
Remote Web Developer Jobs
All Filters
Total selected ()
No Results
No Results






.png)









.png)

















