Top Remote Site Reliability Engineer Jobs

Reposted YesterdaySaved
Remote
USA
136K-237K Annually
Expert/Leader
136K-237K Annually
Expert/Leader
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills: Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
YesterdaySaved
Remote or Hybrid
Location, WV, USA
Junior
Junior
Payments
Supports the availability, performance, and reliability of a global payment platform. Responsibilities include monitoring system health, improving observability, tuning alerts, developing dashboards, investigating incidents, troubleshooting integrations and data-processing issues, and creating automation scripts and AI-powered workflows. The role collaborates with software and infrastructure teams, participates in an on-call rotation, and develops skills in cloud operations, platform engineering, incident response, and responsible AI usage.
Top Skills: AWSAzureAzure Ai FoundryAzure Ai ServicesChatgptCi/CdClaudeDatadogDnsDynatraceGitGCPHttp/HttpsInfrastructure As CodeKubernetesLoad BalancingNew RelicPowershellPythonSplunkSQLTcp/Ip
Reposted YesterdaySaved
Remote
United States
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Reposted YesterdaySaved
In-Office or Remote
4 Locations
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Designs, operates, and improves large-scale Microsoft 365 and Purview services. Responsibilities include developing automation scripts, building telemetry pipelines and monitoring tools, troubleshooting production issues, optimizing code and services, participating in engineering reviews, and responding to incidents through on-call rotations. The role focuses on cloud engineering, service reliability, observability, resiliency, security, and operational excellence for enterprise and government customers.
Top Skills: Ai-Powered Operational ToolingAutomationCC#C++Cloud EngineeringDistributed SystemsJavaJavaScriptMicrosoft 365Microsoft PurviewMonitoring ToolsObservabilityPythonTelemetry Pipelines
Reposted 29 Days AgoSaved
Easy Apply
Remote or Hybrid
USA
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 29 Days AgoSaved
Easy Apply
Remote or Hybrid
Virginia, USA
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
As an intern, manage operational tasks in classified environments, develop automation tools, create documentation, and enhance services for Zscaler's cloud security platform.
Top Skills: Aws EcsKubernetesPython
Reposted YesterdaySaved
In-Office or Remote
Boston, MA, USA
166K-308K Annually
Senior level
166K-308K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Lead design, build, and evolve developer infrastructure and CI platforms for Meraki cloud teams. Guide complex troubleshooting, mentor engineers, drive operational excellence, define roadmaps with leadership, and champion sustainable on-call practices while supporting large-scale distributed systems and automation across developer environments.
Top Skills: Artifact ManagementBare MetalBuild ToolsCiCi PlatformsCode ReviewConfiguration-As-CodeContainer OrchestrationContainerizationInfrastructure AutomationManaged Cloud ServicesPythonRubyUnix/Linux
2 Days AgoSaved
In-Office or Remote
2 Locations
76K-136K Annually
Entry level
76K-136K Annually
Entry level
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills: AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Reposted 2 Days AgoSaved
In-Office or Remote
5 Locations
129K-256K Annually
Senior level
129K-256K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Operate, deploy, and optimize backend collaboration services for global SaaS. Build and evolve CI/CD and automation, lead incident response and RCA, use observability for capacity planning, and define operational best practices and runbooks to improve reliability and scalability across cloud and hybrid environments.
Top Skills: BashCi/CdDockerGitGoInfrastructure-As-CodeKubernetesLinuxMonitoringObservabilityPython
2 Days AgoSaved
In-Office or Remote
2 Locations
135K-160K Annually
Senior level
135K-160K Annually
Senior level
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills: Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
2 Days AgoSaved
Remote
US
120K-170K Annually
Senior level
120K-170K Annually
Senior level
Cloud • Software
Operate and improve large-scale on-premises infrastructure across bare-metal servers, VMs, storage, networking, Kubernetes, containers, CI/CD, and Kafka. Automate provisioning and configuration with Ansible and related infrastructure-as-code tools, maintain observability and reliability, troubleshoot hardware and operating systems, manage disaster recovery, and support incident response. The role requires onsite work weekly at a designated data center and participation in on-call rotations.
Top Skills: AnsibleArgo CdBare-Metal ServersBashCertificate ManagementCi/CdContainersDhcpDnsElk StackFirewallsForemanGitopsGoGrafanaHelmKafkaKubernetesLinuxLinux NetworkingLoad BalancersMaasNtpPrometheusPythonRoutingStorage AppliancesTerraformVirtual MachinesVirtualizationVlans
Reposted One Month AgoSaved
Easy Apply
Remote or Hybrid
5 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills: AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
5 Days AgoSaved
In-Office or Remote
Santa Clara, CA, USA
200K-322K Annually
Senior level
200K-322K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead the transformation, design, deployment, and operation of globally scaled on-premises and cloud compute infrastructure. Build reliable core services including DNS, NTP/PTP, DHCP, and LDAP; develop automation, monitoring, capacity planning, and lifecycle management. Define performance metrics, optimize systems using technologies such as SR-IOV and DPUs, and develop data analysis and visualization tools. Partner with engineering, product, program, and company leadership to deliver scalable IT services.
Top Skills: AnycastBare MetalBgpConfiguration Management ToolsContainersDhcpDnsDpuEbpfGoInfrastructure As CodeLdapLinuxLinux KernelMicroservicesNtpPtpPythonSdnSr-IovTerraformVlanVxlanXdp
5 Days AgoSaved
Remote
United States
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Own reliability and operational health for Microsoft Substrate cloud services in regulated environments. Participate in on-call rotations, diagnose and resolve production incidents, build automation to reduce toil, maintain production code, develop monitoring and telemetry supporting SLOs, lead post-incident reviews, and collaborate with software engineering teams to improve service reliability, scalability, security, and operability.
Top Skills: Exchange OnlineMicrosoft 365 CopilotMicrosoft CloudMicrosoft Substrate
5 Days AgoSaved
Remote
United States
90K-100K Annually
Mid level
90K-100K Annually
Mid level
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills: Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
Reposted 10 Days AgoSaved
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
5 Days AgoSaved
Remote
United States
125K-150K Annually
Mid level
125K-150K Annually
Mid level
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills: ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
6 Days AgoSaved
Remote or Hybrid
USA
136K-181K Annually
Entry level
136K-181K Annually
Entry level
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills: Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
7 Days AgoSaved
In-Office or Remote
Washington, DC, USA
123K-150K Annually
Entry level
123K-150K Annually
Entry level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Operate and scale Kong’s global multi-region SaaS platform across AWS, GCP, and Azure. Build Kubernetes infrastructure, Terraform-based automation, Helm and ArgoCD deployment workflows, CI/CD pipelines, observability systems, and highly available data layers. Improve Kong Gateway and Mesh environments, reliability, scalability, security, disaster recovery, and cost efficiency. Participate in 24/7 on-call, incident response, SLO tracking, postmortems, and operational improvement initiatives.
Top Skills: ArgocdAWSAzureAzure VnetBashCi/CdClickhouseDatadogDnsDruidGCPGcp NccGitopsGoGrafanaHelmHTTPKafkaKong GatewayKong MeshKubernetesLinuxLoad BalancersPostgresPrivatelinkPrometheusPythonRedisTerraformTerragruntThanosTls/SslTransit GatewayVpc Peering
Reposted 7 Days AgoSaved
Remote or Hybrid
3 Locations
174K-238K Annually
Senior level
174K-238K Annually
Senior level
Cloud
The Staff Site Reliability Engineer will manage large-scale cloud production systems, ensuring reliability and performance, while automating processes and responding to incidents.
Top Skills: AWSBashCloudFormationDockerGoHelmKubernetesPythonRubyTerraform
8 Days AgoSaved
Remote
US
174K-305K Annually
Senior level
174K-305K Annually
Senior level
Artificial Intelligence • Software
Architects and owns highly available infrastructure and Kubernetes-based platforms supporting autonomous systems. Builds Golang backend services, platform tooling, observability systems, dashboards, alerts, and log aggregation. Partners with product teams to launch services, performs performance analysis, manages cloud upgrades, and participates in incident response and postmortems. Collaborates on cloud security risk assessments, intrusion detection, threat-feed systems, risk mitigation, and SaaS payment processes. Provides architectural leadership and mentorship across engineering teams.
Top Skills: ArgocdArgocd Image UpdaterArtifactoryAWSGithub ActionsGoJavaScriptKubernetesPythonRustTerraform
Reposted 8 Days AgoSaved
In-Office or Remote
3 Locations
120K-261K Annually
Mid level
120K-261K Annually
Mid level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, and operate highly reliable, secure distributed systems for Azure Data Transfer. Define SLIs/SLOs, reduce toil through automation and IaC, improve observability, lead incident response/on-call, drive progressive delivery and safe rollouts, and ensure compliance with security and audit requirements.
Top Skills: AnsibleAzureAzure DevopsCi/CdGithub ActionsInfrastructure As CodeMarinerRed HatRocky 9
Reposted 13 Days AgoSaved
Easy Apply
Remote or Hybrid
9 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 8 Days AgoSaved
In-Office or Remote
9 Locations
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
9 Days AgoSaved
In-Office or Remote
2 Locations
192K-356K Annually
Expert/Leader
192K-356K Annually
Expert/Leader
Cloud • Information Technology • Internet of Things • Professional Services • Software
Leads technical strategy for site reliability, critical incident response, and production resilience across Splunk Cloud. Serves as the escalation authority during P1/P2 events, owns complex enterprise customer environments, guides infrastructure architecture and automation, leads post-mortems and root-cause analysis, improves observability and operational processes, and mentors senior engineers without formal management authority.
Top Skills: AWSGoGoogle Cloud Platform (Gcp)Indexer ClusteringKvstoreLinuxAzurePythonSearch Head ClustersSplSplunkSplunk Observability
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account