Maximum of 25 job preferences reached.
Top Remote Site Reliability Engineer Jobs
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills:
BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, resilience, security, and performance optimization of enterprise mainframe environments. Responsibilities include z/OS performance tuning, WLM and RACF administration, business continuity planning, automation, technical governance, incident resolution, stakeholder collaboration, and guidance of cross-functional engineering and operations teams. The role also evaluates cloud, DevOps, AI, and hybrid IT technologies for mainframe transformation.
Top Skills:
AnsibleCsmGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRed Hat Ansible For Ibm Z CollectionsRmfSmfWlmZlinux
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills:
AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Deploy, monitor, automate, and support large-scale distributed IaaS, PaaS, and SaaS environments. Build reliable infrastructure, measure production performance, resolve complex service issues, scale systems through automation, and collaborate across engineering, DevOps, security, and IT operations teams. The role requires expertise in networking, cloud technologies, storage, virtualization, infrastructure automation, and service lifecycle management, with eligibility for Secret and TS/SCI clearances.
Top Skills:
AnsibleAzure StackCephHelmIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
2 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills:
Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
Reposted 3 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
Financial Services
Lead site reliability engineering for enterprise infrastructure platforms. Build automated, self-healing systems; define and operationalize SLIs, SLOs, observability, and actionable alerting; own production services, incidents, postmortems, security, performance, and cost. Mentor engineers, lead resiliency reviews, reduce toil, and apply validated enterprise AI capabilities to incident response, SDLC workflows, and operational readiness while maintaining security, traceability, and reliability controls.
Top Skills:
C++Ci/CdDatadogDynatraceGoGrafanaJavaKubernetesPrometheusPythonRustSplunkTerraform
Artificial Intelligence • Machine Learning • Software • Defense
Own platform reliability, observability, incident response, scaling, capacity planning, and deployment automation. Monitor system health, debug and resolve incidents, build logging and monitoring tools, improve CI/CD pipelines, develop self-service automation, and strengthen high-availability delivery systems. The role requires clear incident communication, post-incident learning, and collaboration across engineering teams in secure, high-side environments.
Top Skills:
AWSAws GovcloudBashCi/CdDatadogDockerElasticsearchOpensearchPlg StackPulumiPythonSQLTerraform
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills:
AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Financial Services
Leads reliability engineering for mission-critical network services, including resiliency reviews, incident response, root-cause analysis, automation, observability, and durable remediation. Architects self-healing and guarded remediation workflows using Python, Shell, and Ansible. Provides technical leadership across SD-WAN, SDN, routing, switching, security, and traffic services while embedding SRE practices, resilience testing, AI-assisted workflows, and security controls. Mentors engineers and drives service-level objectives, error budgets, and operational readiness.
Top Skills:
.NetAnsibleCi/CdContainer OrchestrationContainersFirewallsJavaLoad BalancersObservabilityProxiesPythonRoutingSd-WanSdaSdnShellSpring BootSwitching
17 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Architect, operate, and maintain resilient cloud infrastructure across Azure and AWS commercial and government environments. Support Kubernetes platforms, IaC, observability, monitoring, deployments, platform services, performance testing, and incident response. Define reliability metrics, participate in 24/7 on-call rotations, perform root cause analysis, and automate operational processes and remediation. The role requires U.S. citizenship and eligibility to obtain a Confidential, Secret, or Top Secret clearance.
Top Skills:
ArgocdAWSAzureAzure MonitorDynatraceEncryptionFluxGitGitlabGitopsGrafanaHelmIaasIamKubernetesOwaspPaasPkiPrometheusPulumiRestful ApisSplunkTerraformVisual Studio Code
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills:
AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Other • Retail
Lead and develop an SRE team responsible for the reliability, availability, performance, automation, and observability of Linux-based digital commerce infrastructure. Set SRE and DevOps strategy, modernize Kubernetes and CI/CD practices, establish automation and infrastructure-as-code standards, oversee incident response, and drive SLIs, SLOs, and error budgets. Partner across engineering, architecture, infrastructure, security, networking, and product teams while recruiting, mentoring, and developing SRE talent.
Top Skills:
Apache TomcatCi/CdDatadogDockerError BudgetsF5Github ActionsInfrastructure As CodeJfrog ArtifactoryKubernetesLinuxNginxPuppetPythonSlis/SlosTerraformVMware
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills:
AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
Information Technology • Consulting • Defense
Operate and improve reliable Azure Government platforms supporting AI applications, OIGChat, and enterprise data systems. Responsibilities include monitoring, observability, incident response, root-cause analysis, performance and cost optimization, capacity planning, disaster recovery, high availability, deployment readiness, and operational documentation. The role also supports Azure Databricks, data pipelines, AI model endpoints, secure government cloud environments, and mission-critical federal operations.
Top Skills:
Ai/Ml PlatformsApplication InsightsAws GovcloudAzureAzure DatabricksAzure MonitorAzure OpenaiAzure SynapseDevOpsElk StackGrafanaLog AnalyticsPaasPrometheusSre
Artificial Intelligence • Machine Learning • Software • Analytics
Leads and builds Infinia’s Release Engineering, DevOps, SecDevOps, and SRE functions. Owns release pipelines, CI/CD automation, infrastructure as code, security tooling, reliability standards, observability, incident response, capacity planning, and scaling. Establishes SLOs, SLIs, error budgets, and operational practices while building globally distributed engineering teams. Partners with product, QA, security, and executive leadership to communicate roadmaps, risks, and technical tradeoffs.
Top Skills:
AnsibleCi/CdDastDependency ScanningDockerGithub ActionsGitlab CiGrafanaJenkinsKubernetesOpentelemetryPrometheusPulumiSastSecrets ManagementTektonTerraform
Software
Lead reliability, scalability, observability, and incident-management initiatives for critical distributed systems. Define SLOs, error budgets, and actionable alerts; automate toil; improve production readiness and recovery; lead incident response and postmortems; build reusable operational tooling; partner with engineering and product teams on architecture; and mentor engineers while scaling SRE practices.
Top Skills:
AlertingAWSAzureError BudgetsGCPGoInfrastructure As CodeJavaKubernetesLoggingMetricsObservabilityPythonSlisSlosTerraformTracing
Edtech • Fintech • Information Technology • Software
Operate and improve AWS production infrastructure, infrastructure as code, Kubernetes workloads, observability, CI/CD, and incident response. The role investigates root causes, strengthens application resilience, automates operational tasks, supports database and performance reliability, maintains documentation, and participates in 24/7 on-call rotations. The engineer owns scoped reliability projects and collaborates with product engineering teams on resilient, secure, and compliant systems.
Top Skills:
Amazon EksAmazon RdsAWSCircleCIDatadogGithub ActionsKubernetesLinuxNew RelicOpensearchPostgresRedisRubyRuby On RailsTerraform
Hardware • Healthtech
Owns the reliability, security, performance, and availability of AWS-hosted healthcare infrastructure. Builds infrastructure as code, CI/CD automation, monitoring, observability, backup and disaster recovery capabilities. Leads incident response, optimizes cloud resources, implements security controls, supports customer onboarding and migrations, and ensures compliance with healthcare privacy and software lifecycle requirements. Participates in on-call rotations and provides technical guidance to engineering and support teams.
Top Skills:
AWSBashCi/CdCitrixEcsGdprHipaaHyper-VIec 62304JavaScriptJinjaJSONMirth ConnectPythonTerraformTypescriptVMwareYaml
Information Technology • Consulting
Administer and secure the organization’s GitHub environment, including repositories, permissions, branch protections, security controls, and CI/CD workflows. Build automation with GitHub Actions, APIs, and scripting; monitor reliability against SLOs; troubleshoot incidents; and improve developer experience. Integrate identity providers and security tools, support migrations, maintain documentation, and guide teams on GitHub usage and Copilot adoption. The role requires SRE or DevOps experience, infrastructure-as-code, containers, cloud platforms, and observability tooling.
Top Skills:
AnsibleAWSAzureBashCodeqlDatadogDependabotDockerGCPGithub ActionsGithub ApiGithub CliGithub CopilotGithub EnterpriseGrafanaPowershellPrometheusPythonSAMLScimSplunkSsoTerraform
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills:
AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills:
AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Remote Site Reliability Engineers
See AllPopular Job Searches
All Remote Software Engineer Jobs
Remote .NET Developer Jobs
Remote AI Engineer Jobs
Remote Android Developer Jobs
Remote Android Engineer Jobs
Remote Automation Engineer Jobs
Remote AWS Jobs
Remote Backend Engineer Jobs
Remote C# Jobs
Remote C++ Jobs
Remote Cloud Architect Jobs
Remote Cloud Engineer Jobs
Remote Design Engineer Jobs
Remote DevOps Engineer Jobs
Remote DevOps Jobs
Remote Embedded Software Engineer Jobs
Remote Engineering Director Jobs
Remote Engineering Manager Jobs
Remote Enterprise Architect Jobs
Remote Field Engineer Jobs
Remote Front-End Developer Jobs
Remote Front-End Engineer Jobs
Remote Full-Stack Engineer Jobs
Remote Game Developer Jobs
Remote Golang Jobs
Remote Hardware Engineer Jobs
Remote Infrastructure Engineer Jobs
Remote Integration Engineer Jobs
Remote iOS Developer Jobs
Remote iOS Engineer Jobs
Remote IT Engineer Jobs
Remote Java Developer Jobs
Remote Javascript Jobs
Remote Lead Software Engineer Jobs
Remote Linux Engineer Jobs
Remote Linux Jobs
Remote Network Engineer Jobs
Remote Perl Jobs
Remote PHP Developer Jobs
Remote Platform Engineer Jobs
Remote Principal Software Engineer Jobs
Remote Project Engineer Jobs
Remote Python Developer + Engineer Jobs
Remote Python Jobs
Remote QA Analyst Jobs
Remote QA Automation Engineer Jobs
Remote QA Engineer Jobs
Remote Ruby Jobs
Remote Sales Engineer Jobs
Remote Salesforce Administrator Jobs
Remote Salesforce Developer Jobs
Remote Salesforce Developer Jobs
Remote Scala Jobs
Remote Senior DevOps Engineer Jobs
Remote Software Architect Jobs
Remote Software Development Manager Jobs
Remote Software Engineering Manager Jobs
Remote Solutions Architect Jobs
Remote Solutions Engineer Jobs
Remote SRE Jobs
Remote Staff Software Engineer Jobs
Remote Systems Engineer Jobs
Remote Tech Lead Jobs
Remote Test Engineer Jobs
Remote VP of Engineering Jobs
Remote Web Developer Jobs
All Filters
Total selected ()
No Results
No Results




















.jpg)





.jpg)




