Top Remote Site Reliability Engineer Jobs

Reposted 10 Days AgoSaved
Remote
USA
Senior level
Senior level
Information Technology • Cryptocurrency
The Site Reliability Engineer will lead technical initiatives, architect solutions, troubleshoot issues, mentor team members, and improve observability practices.
Top Skills: ArgocdBashElk StackGCPGoGrafanaHelmKubernetesPrometheusPythonTerraform
Reposted 10 Days AgoSaved
Remote
United States
154K-231K Annually
Senior level
154K-231K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Lead technical design and architecture of large-scale Ceph clusters, drive capacity planning and major upgrades, resolve complex cross-functional production incidents, establish automation and reliability standards, and mentor engineers on advanced Ceph operations to scale the global storage platform.
Top Skills: AnsibleBluestoreCephCephfsCinderCrushCsiDdnErasure CodingGoIsilonKubernetesManilaMdsMgrMonNetappNeutronNovaOpenstackOsdPlacement GroupsPurePythonRbdRbd MirroringRgwRgw MultisiteRookS3SaltstackSwiftTerraformVastWeka
Reposted 10 Days AgoSaved
Remote
United States
128K-192K Annually
Senior level
128K-192K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Own and maintain reliability, performance, scalability, and capacity of massive Ceph storage clusters. Troubleshoot distributed storage issues, build automation (Python/Shell, SaltStack, Ansible), define observability (SLIs/SLOs, PromQL/LogQL, Grafana), and lead lifecycle initiatives like upgrades, expansions, and migrations.
Top Skills: AnsibleCeph (RadosCephfs)ChefGrafanaLinuxLogqlPromqlPuppetPythonRbdRgwSaltstackShell
Reposted 11 Days AgoSaved
Remote
United States
100K-140K Annually
Mid level
100K-140K Annually
Mid level
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills: DhcpDnsLinuxNtpPython
12 Days AgoSaved
In-Office or Remote
3 Locations
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Embed with customer teams running large-scale GPU training and inference to onboard, tune, debug, and improve reliability. Diagnose fabric, driver, scheduler, and application failures; profile performance; build automation, monitoring, and preflight checks; lead incident response; and convert field learnings into product improvements and reusable reference configurations.
Top Skills: AnsibleBashCgroupsContainer RuntimesCudaDcgmDevice PluginsFabric ManagerGoGpfsHelmInfinibandKubernetesKv CacheLustreNamespacesNcclNvidia DriversNvidia-SmiNvlinkPythonRoceSlurmSshTerraformTopology-Aware SchedulingVastWeka
12 Days AgoSaved
In-Office or Remote
Jupiter, FL, USA
Mid level
Mid level
Cloud • Other
Maintain and monitor Rocket.net hosting platform reliability and performance. Provide advanced escalation-level technical support for WordPress VIP customers, troubleshoot Linux-based production environments, web servers, databases, caching, DNS, and networking. Participate in incident response, root cause analysis, automation, documentation, and collaborate with support and engineering teams to improve platform stability.
Top Skills: ApacheBashCachingCdnCloudflareDatadogDnfDnsHTTPHttpsLinuxMariadbMySQLNedataNginxPhp-FpmRedisSshSsl/TlsWafWordpressYum
12 Days AgoSaved
Remote
USA
75K-90K Annually
Mid level
75K-90K Annually
Mid level
3PL: Third Party Logistics
Own and improve uptime for backend services, APIs, workers, and ML pipelines; build monitoring, alerting, auto-remediation, optimize GCP infrastructure and Postgres performance, run on-call rotation and blameless postmortems.
Top Skills: Apollo ServerBashCi/CdCloud RunCloud SqlDatadogDockerExpressGCPGcsGrafanaNext.JsNode.jsPostgresPrometheusPythonReactTerraformTypescriptZabbix
12 Days AgoSaved
Remote
US
175K-185K Annually
Senior level
175K-185K Annually
Senior level
Software
Lead improvements in reliability, performance, scalability, capacity, and observability through automation and tooling. Partner with developers to design infrastructure and monitoring, define SLIs/SLOs, conduct load and performance testing, participate in incident response and root cause analysis, and manage monitoring services for production systems.
Top Skills: DockerGoGradleJavaKubernetesOpentelemetrySpring BootTerraform
Reposted 12 Days AgoSaved
In-Office or Remote
17 Locations
Senior level
Senior level
Fintech • Information Technology • Software • Financial Services
Design, build, and maintain real-time, secure distributed systems and observability UIs/APIs. Implement CI/CD, containerized deployments (Docker/Kubernetes/OpenShift), integrate observability stack (Elasticsearch/Logstash/Grafana), and apply secure coding and API security standards to ensure reliability, performance, and incident automation. Collaborate in Agile teams and explore AI to improve resiliency.
Top Skills: Agentic AiCi/CdDockerElasticsearchGrafanaJava Spring BootKafkaKubernetesLogstashMariadbNode.jsOauth2OpenshiftReactSecrets Management
Reposted 12 Days AgoSaved
In-Office or Remote
9 Locations
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
13 Days AgoSaved
In-Office or Remote
12 Locations
140K-150K Annually
Mid level
140K-150K Annually
Mid level
Healthtech
Build, operate, and scale AWS cloud infrastructure and Kubernetes workloads using Terraform and Helm. Improve observability, define SLIs/SLOs, automate deployments and incident response, support on-call rotation, and implement security and compliance (HIPAA, SOC 2) best practices while partnering with product and engineering teams.
Top Skills: AWSCi/CdEvent SourcingHelmKubernetesLinuxMonitoring/Logging/TracingNetworkingTerraform
13 Days AgoSaved
In-Office or Remote
6 Locations
132K-176K Annually
Junior
132K-176K Annually
Junior
Security • Software
Build, scale, and maintain Tenable's cloud-based vulnerability management platform for private and U.S. government customers. Troubleshoot escalations, automate deployments and monitoring, collaborate across engineering teams, document operational procedures, participate in on-call rotation, remediate infrastructure/container security issues, and contribute to reliability, standardization, and production deployments.
Top Skills: BashContainersDockerKubernetesMicroservicesNode.jsPythonTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 18 Days AgoSaved
Easy Apply
Remote or Hybrid
9 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
14 Days AgoSaved
Remote
USA
140K-165K Annually
Senior level
140K-165K Annually
Senior level
Hardware • Machine Learning • Security • Software
Design and maintain developer experience tooling and CI/CD pipelines, manage self-service deployment tools (secrets, rollbacks), partner with cloud teams to troubleshoot production systems, participate in on-call rotation, and write production-grade automation and tests in Go and TypeScript to improve developer velocity and reliability.
Top Skills: AWSGithub ActionsGoHelmKubernetesTerraformTypescript
Reposted 14 Days AgoSaved
In-Office or Remote
3 Locations
120K-261K Annually
Mid level
120K-261K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills: AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Reposted 14 Days AgoSaved
Remote
United States
115K-135K Annually
Mid level
115K-135K Annually
Mid level
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills: ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
15 Days AgoSaved
In-Office or Remote
Location, WV, USA
Mid level
Mid level
Healthtech • Telehealth
SRE responsible for designing and maintaining observability across Azure and multi-cloud environments, defining SLIs/SLOs, building dashboards and automated alerts, running incident response and on-call playbooks, driving reliability automation and self-healing, contributing to disaster recovery, and partnering with security and engineering teams to enforce governance and best practices.
Top Skills: AksAnsibleApplication InsightsAzureAzure MonitorBicepChaos MeshDatadogDynatraceElasticGrafanaGremlinKubernetesLogicmonitorPowershellPrometheusPythonTerraformVms
20 Days AgoSaved
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills: .NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Reposted 15 Days AgoSaved
Remote
United States
Mid level
Mid level
Blockchain • Software
Build, operate, and scale production Kubernetes infrastructure using GitOps and declarative IaC. Design CI/CD workflows, observability, and secure-by-default systems. Troubleshoot networking/storage, participate in on-call rotations, automate operational workflows, and drive postmortems and reliability improvements.
Top Skills: ArbitrumArgocdArgocd ApplicationsetsAWSAzureBashCloudwatchCodebuildGCPGithub ActionsGitopsGoGrafanaK9SKubernetesLinuxLokiMimirPrometheusPrysmPythonTerraformYamlZerodev
16 Days AgoSaved
In-Office or Remote
4 Locations
110K-120K Annually
Mid level
110K-120K Annually
Mid level
Fintech • Software
Operate and improve the reliability, availability, and security of a FedRAMP High cloud platform as an SRE/L3 escalation point. Monitor systems, troubleshoot incidents, lead incident response and RCA, build runbooks, enhance observability and automation, support deployments, and ensure compliance while partnering with engineering teams to reduce toil and improve resilience.
Top Skills: AWSAws BackupBashCi/CdCloudwatchConfiguration ManagementDatadogEksFedramp HighGoGrafanaIamInfrastructure As CodeIstioJira Service ManagementKubernetesLinuxOpentelemetryPagerdutyPowershellPrometheusPythonRdsRoute 53SplunkVpc
16 Days AgoSaved
In-Office or Remote
3 Locations
110K-130K Annually
Mid level
110K-130K Annually
Mid level
Fintech • Software
Owner of operational health and availability for a FedRAMP High cloud platform. Act as L3 escalation, monitor observability, lead incident response, perform root cause analysis, automate toil, support deployments, maintain runbooks, and ensure compliance while partnering with engineering to improve reliability and resilience.
Top Skills: AWSAws BackupBashCi/CdCloudwatchDatadogEksGoGrafanaIamIstioJira Service ManagementKubernetesLinuxOpentelemetryPagerdutyPowershellPrometheusPythonRdsRoute 53SplunkVpc
Reposted 16 Days AgoSaved
Remote
USA
Senior level
Senior level
Software • Web3
Lead reliability practices across teams: embed early in projects, define SLIs/SLOs, build multi-cloud paved roads with Terraform, run on-call, drive org-wide incident maturity and tooling.
Top Skills: AWSAzureGCPRuby On RailsTerraformTypescriptWebcontainers
Reposted 16 Days AgoSaved
In-Office or Remote
2 Locations
165K-215K Annually
Senior level
165K-215K Annually
Senior level
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills: AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Reposted 16 Days AgoSaved
In-Office or Remote
7 Locations
200K-200K Annually
Mid level
200K-200K Annually
Mid level
Cloud • Software
The Site Reliability Engineer will ensure reliable cloud operations by applying Python for infrastructure automation, managing OpenStack and Kubernetes, and practicing devsecops in a fast-paced environment.
Top Skills: KubernetesLinuxOpenstackPython
Reposted 16 Days AgoSaved
In-Office or Remote
2 Locations
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills: GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account