Top Remote Site Reliability Engineer Jobs

YesterdaySaved
In-Office or Remote
Eden Prairie, MN, USA
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills: BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
YesterdaySaved
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, resilience, security, and performance optimization of enterprise mainframe environments. Responsibilities include z/OS performance tuning, WLM and RACF administration, business continuity planning, automation, technical governance, incident resolution, stakeholder collaboration, and guidance of cross-functional engineering and operations teams. The role also evaluates cloud, DevOps, AI, and hybrid IT technologies for mainframe transformation.
Top Skills: AnsibleCsmGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRed Hat Ansible For Ibm Z CollectionsRmfSmfWlmZlinux
YesterdaySaved
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
2 Days AgoSaved
Remote or Hybrid
Merrimack, NH, USA
118K-201K Annually
Senior level
118K-201K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Deploy, monitor, automate, and support large-scale distributed IaaS, PaaS, and SaaS environments. Build reliable infrastructure, measure production performance, resolve complex service issues, scale systems through automation, and collaborate across engineering, DevOps, security, and IT operations teams. The role requires expertise in networking, cloud technologies, storage, virtualization, infrastructure automation, and service lifecycle management, with eligibility for Secret and TS/SCI clearances.
Top Skills: AnsibleAzure StackCephHelmIaasJdfsJuniperKubernetesNfsOpenstackPaasS3SaaSSecurity+TerraformVMware
2 Days AgoSaved
Easy Apply
Remote
3 Locations
Easy Apply
223K-380K Annually
Expert/Leader
223K-380K Annually
Expert/Leader
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills: Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
Reposted 3 Days AgoSaved
Easy Apply
Remote or Hybrid
9 Locations
Easy Apply
119K-170K Annually
Senior level
119K-170K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills: BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
6 Days AgoSaved
Remote or Hybrid
2 Locations
Senior level
Senior level
Financial Services
Lead site reliability engineering for enterprise infrastructure platforms. Build automated, self-healing systems; define and operationalize SLIs, SLOs, observability, and actionable alerting; own production services, incidents, postmortems, security, performance, and cost. Mentor engineers, lead resiliency reviews, reduce toil, and apply validated enterprise AI capabilities to incident response, SDLC workflows, and operational readiness while maintaining security, traceability, and reliability controls.
Top Skills: C++Ci/CdDatadogDynatraceGoGrafanaJavaKubernetesPrometheusPythonRustSplunkTerraform
9 Days AgoSaved
In-Office or Remote
San Diego, CA, USA
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software • Defense
Own platform reliability, observability, incident response, scaling, capacity planning, and deployment automation. Monitor system health, debug and resolve incidents, build logging and monitoring tools, improve CI/CD pipelines, develop self-service automation, and strengthen high-availability delivery systems. The role requires clear incident communication, post-incident learning, and collaboration across engineering teams in secure, high-side environments.
Top Skills: AWSAws GovcloudBashCi/CdDatadogDockerElasticsearchOpensearchPlg StackPulumiPythonSQLTerraform
11 Days AgoSaved
Remote
USA
170K-225K Annually
Mid level
170K-225K Annually
Mid level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills: AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
Reposted 15 Days AgoSaved
Remote or Hybrid
United States
168K-210K Annually
Senior level
168K-210K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills: AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
16 Days AgoSaved
Remote or Hybrid
2 Locations
Senior level
Senior level
Financial Services
Leads reliability engineering for mission-critical network services, including resiliency reviews, incident response, root-cause analysis, automation, observability, and durable remediation. Architects self-healing and guarded remediation workflows using Python, Shell, and Ansible. Provides technical leadership across SD-WAN, SDN, routing, switching, security, and traffic services while embedding SRE practices, resilience testing, AI-assisted workflows, and security controls. Mentors engineers and drives service-level objectives, error budgets, and operational readiness.
Top Skills: .NetAnsibleCi/CdContainer OrchestrationContainersFirewallsJavaLoad BalancersObservabilityProxiesPythonRoutingSd-WanSdaSdnShellSpring BootSwitching
17 Days AgoSaved
Easy Apply
Remote
United States
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
23 Days AgoSaved
In-Office or Remote
Eden Prairie, MN, USA
73K-130K Annually
Mid level
73K-130K Annually
Mid level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Architect, operate, and maintain resilient cloud infrastructure across Azure and AWS commercial and government environments. Support Kubernetes platforms, IaC, observability, monitoring, deployments, platform services, performance testing, and incident response. Define reliability metrics, participate in 24/7 on-call rotations, perform root cause analysis, and automate operational processes and remediation. The role requires U.S. citizenship and eligibility to obtain a Confidential, Secret, or Top Secret clearance.
Top Skills: ArgocdAWSAzureAzure MonitorDynatraceEncryptionFluxGitGitlabGitopsGrafanaHelmIaasIamKubernetesOwaspPaasPkiPrometheusPulumiRestful ApisSplunkTerraformVisual Studio Code
Reposted 7 Hours AgoSaved
Remote
United States
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Reposted 10 Hours AgoSaved
In-Office or Remote
3 Locations
120K-261K Annually
Mid level
120K-261K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills: AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
YesterdaySaved
Remote
USA
9K-17K Annually
Senior level
9K-17K Annually
Senior level
Other • Retail
Lead and develop an SRE team responsible for the reliability, availability, performance, automation, and observability of Linux-based digital commerce infrastructure. Set SRE and DevOps strategy, modernize Kubernetes and CI/CD practices, establish automation and infrastructure-as-code standards, oversee incident response, and drive SLIs, SLOs, and error budgets. Partner across engineering, architecture, infrastructure, security, networking, and product teams while recruiting, mentoring, and developing SRE talent.
Top Skills: Apache TomcatCi/CdDatadogDockerError BudgetsF5Github ActionsInfrastructure As CodeJfrog ArtifactoryKubernetesLinuxNginxPuppetPythonSlis/SlosTerraformVMware
29 Days AgoSaved
In-Office or Remote
Eden Prairie, MN, USA
135K-231K Annually
Expert/Leader
135K-231K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills: AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
2 Days AgoSaved
In-Office or Remote
Washington, DC, USA
115K-252K Annually
Expert/Leader
115K-252K Annually
Expert/Leader
Information Technology • Consulting • Defense
Operate and improve reliable Azure Government platforms supporting AI applications, OIGChat, and enterprise data systems. Responsibilities include monitoring, observability, incident response, root-cause analysis, performance and cost optimization, capacity planning, disaster recovery, high availability, deployment readiness, and operational documentation. The role also supports Azure Databricks, data pipelines, AI model endpoints, secure government cloud environments, and mission-critical federal operations.
Top Skills: Ai/Ml PlatformsApplication InsightsAws GovcloudAzureAzure DatabricksAzure MonitorAzure OpenaiAzure SynapseDevOpsElk StackGrafanaLog AnalyticsPaasPrometheusSre
2 Days AgoSaved
Remote or Hybrid
North Carolina, USA
250K-300K Annually
Expert/Leader
250K-300K Annually
Expert/Leader
Artificial Intelligence • Machine Learning • Software • Analytics
Leads and builds Infinia’s Release Engineering, DevOps, SecDevOps, and SRE functions. Owns release pipelines, CI/CD automation, infrastructure as code, security tooling, reliability standards, observability, incident response, capacity planning, and scaling. Establishes SLOs, SLIs, error budgets, and operational practices while building globally distributed engineering teams. Partners with product, QA, security, and executive leadership to communicate roadmaps, risks, and technical tradeoffs.
Top Skills: AnsibleCi/CdDastDependency ScanningDockerGithub ActionsGitlab CiGrafanaJenkinsKubernetesOpentelemetryPrometheusPulumiSastSecrets ManagementTektonTerraform
2 Days AgoSaved
Easy Apply
Remote
US
Easy Apply
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Software
Lead reliability, scalability, observability, and incident-management initiatives for critical distributed systems. Define SLOs, error budgets, and actionable alerts; automate toil; improve production readiness and recovery; lead incident response and postmortems; build reusable operational tooling; partner with engineering and product teams on architecture; and mentor engineers while scaling SRE practices.
Top Skills: AlertingAWSAzureError BudgetsGCPGoInfrastructure As CodeJavaKubernetesLoggingMetricsObservabilityPythonSlisSlosTerraformTracing
4 Days AgoSaved
Remote
United States
Senior level
Senior level
Edtech • Fintech • Information Technology • Software
Operate and improve AWS production infrastructure, infrastructure as code, Kubernetes workloads, observability, CI/CD, and incident response. The role investigates root causes, strengthens application resilience, automates operational tasks, supports database and performance reliability, maintains documentation, and participates in 24/7 on-call rotations. The engineer owns scoped reliability projects and collaborates with product engineering teams on resilient, secure, and compliant systems.
Top Skills: Amazon EksAmazon RdsAWSCircleCIDatadogGithub ActionsKubernetesLinuxNew RelicOpensearchPostgresRedisRubyRuby On RailsTerraform
4 Days AgoSaved
Remote
USA
120K-130K Annually
Senior level
120K-130K Annually
Senior level
Hardware • Healthtech
Owns the reliability, security, performance, and availability of AWS-hosted healthcare infrastructure. Builds infrastructure as code, CI/CD automation, monitoring, observability, backup and disaster recovery capabilities. Leads incident response, optimizes cloud resources, implements security controls, supports customer onboarding and migrations, and ensures compliance with healthcare privacy and software lifecycle requirements. Participates in on-call rotations and provides technical guidance to engineering and support teams.
Top Skills: AWSBashCi/CdCitrixEcsGdprHipaaHyper-VIec 62304JavaScriptJinjaJSONMirth ConnectPythonTerraformTypescriptVMwareYaml
4 Days AgoSaved
Remote
United States
120K-140K Annually
Mid level
120K-140K Annually
Mid level
Information Technology • Consulting
Administer and secure the organization’s GitHub environment, including repositories, permissions, branch protections, security controls, and CI/CD workflows. Build automation with GitHub Actions, APIs, and scripting; monitor reliability against SLOs; troubleshoot incidents; and improve developer experience. Integrate identity providers and security tools, support migrations, maintain documentation, and guide teams on GitHub usage and Copilot adoption. The role requires SRE or DevOps experience, infrastructure-as-code, containers, cloud platforms, and observability tooling.
Top Skills: AnsibleAWSAzureBashCodeqlDatadogDependabotDockerGCPGithub ActionsGithub ApiGithub CliGithub CopilotGithub EnterpriseGrafanaPowershellPrometheusPythonSAMLScimSplunkSsoTerraform
9 Days AgoSaved
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
10 Days AgoSaved
Easy Apply
Remote or Hybrid
USA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account