Top Remote Site Reliability Engineer Jobs

6 Days AgoSaved
In-Office or Remote
Eden Prairie, MN, USA
135K-231K Annually
Expert/Leader
135K-231K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills: AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
Reposted 11 Days AgoSaved
Easy Apply
Remote or Hybrid
10 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
13 Days AgoSaved
Easy Apply
Remote
USA
Easy Apply
241K-270K Annually
Senior level
241K-270K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills: AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
Reposted 13 Days AgoSaved
Remote or Hybrid
4 Locations
175K-175K Annually
Senior level
175K-175K Annually
Senior level
eCommerce • Legal Tech • Professional Services • Software • Data Privacy
The Site Reliability Engineer will ensure systems run smoothly, work with automation tools, resolve issues, and drive operational improvements.
Top Skills: AWSAzureCloudFormationDockerGCPGrafanaKubernetesMemcachedNew RelicOpentelemetryPostgresPrometheusPulumiRedisSentryTerraform
Reposted 15 Days AgoSaved
In-Office or Remote
4 Locations
105K-300K Annually
Entry level
105K-300K Annually
Entry level
Information Technology • Software • Financial Services • Big Data Analytics
SREs at Citadel focus on optimizing and maintaining system reliability, performance, and automation for investment applications, collaborating closely with teams.
Top Skills: Ci/CdCSSJavaScriptPythonReactSQL
Reposted 15 Days AgoSaved
Easy Apply
Remote or Hybrid
Crystal City, VA, USA
Easy Apply
140K-200K Annually
Senior level
140K-200K Annually
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Responsible for managing operations within classified environments, overseeing cloud infrastructure, automating tasks, and ensuring system stability in a high-security setting.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
16 Days AgoSaved
Remote
US
125K-174K Annually
Expert/Leader
125K-174K Annually
Expert/Leader
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills: Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform
Reposted 17 Days AgoSaved
Easy Apply
Remote
United States
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Reposted 17 Days AgoSaved
Remote or Hybrid
Orlando, FL, USA
Expert/Leader
Expert/Leader
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
The Staff Site Reliability Engineer is responsible for ensuring the reliability, performance, and security of workplace collaboration services, focusing on automation, incident management, and operational excellence while providing technical leadership and mentoring to engineers.
Top Skills: Ai EngineeringAzure Virtual DesktopDefender For Office 365Exchange OnlineGraph ApiIntuneJamf ProMicrosoft 365Microsoft Entra IdMicrosoft PurviewOnedrivePowershellSharepoint OnlineTeams
18 Days AgoSaved
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Reposted 19 Days AgoSaved
In-Office or Remote
New York, NY, USA
150K-250K Annually
Mid level
150K-250K Annually
Mid level
Mobile • Software
Site Reliability Engineers will work on production infrastructure, focusing on AWS and Kubernetes while ensuring high availability and customer satisfaction.
Top Skills: AirflowAWSCircleCICloudwatchEksGrafanaMongoDBPagerdutyPingdomRustScala SparkTerraformTypescript
Reposted 19 Days AgoSaved
Easy Apply
Remote or Hybrid
6 Locations
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
21 Days AgoSaved
Remote or Hybrid
New York, NY, USA
Senior level
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead SRE for cloud-based live linear playout systems driving reliability, observability, incident response, SLIs/SLOs, automation, capacity planning, runbooks, and L1/L2 on-call support to ensure resilient distribution across NBCUniversal channels.
Top Skills: AmagiAWSCmafDockerEsamGrafanaH.264HarrisHevcHlsImagineIp NetworkingKubernetesLinuxMicrosoft TeamsRistScte-224Scte-35ServicenowSlackSnellSplunkSrtTs
22 Days AgoSaved
In-Office or Remote
Basking Ridge, NJ, USA
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead enterprise SRE, DevOps, ITSM, and operational excellence for critical banking platforms. Drive reliability (SLI/SLO), automation, CI/CD, IaC, observability, incident/problem/change management, disaster recovery, and AI-enabled operational improvements while building and mentoring cross-functional teams.
Top Skills: AiopsAWSAzureChatopsCi/CdDatadogGitopsGrafanaInfrastructure-As-CodeKubernetesLlmObservabilityOn-PremOpenshiftOpentelemetryPrometheusSplunk
Reposted 23 Days AgoSaved
Easy Apply
Remote or Hybrid
Ontario, CA, USA
Easy Apply
Senior level
Senior level
Artificial Intelligence • Marketing Tech • Software
Lead technical reliability initiatives across a multi-cloud, multi-region active-active content platform. Architect and evolve core services, observability and logging, automation and capacity planning. Mentor engineers, drive cross-team reliability projects, define standards (IaC, SLOs, on-call) and proactively improve platform scalability and incident outcomes.
Top Skills: Apache KafkaApache PulsarAWSCassandraChefEksGCPGkeGoGrafana AlloyGrafana LokiKubernetesLinuxNode.jsPrometheusPythonRubyScylladbShell ScriptingTempoTerraformThanos
YesterdaySaved
In-Office or Remote
Eden Prairie, MN, USA
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Build and operate AWS platforms, improve reliability and scalability, define SLOs, SLAs, and SLIs, automate provisioning and remediation, integrate monitoring and logging, support incident response and on-call operations, perform root cause analysis and postmortems, and develop observability tooling for engineering teams.
Top Skills: AWSAws CdkAws CloudformationCloudwatchDynatraceEc2EcsEksFedramp ModerateGitGitlabIamLambdaLinuxNist 800-171PowershellPythonS3Shell ScriptingTerraformVpcVpn
Reposted 25 Days AgoSaved
In-Office or Remote
Minnetonka, MN, USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Define and scale SRE standards across teams, implement SLOs/SLIs/error budgets, build observability and resiliency patterns, drive automation and AIOps, improve reliability for large-scale Azure cloud systems, and influence engineering and platform teams.
Top Skills: Ai/MlAiopsAutomationAzureError BudgetsIncident ManagementLogsObservability (MetricsOpentelemetrySlisSlosTracing)
Reposted 25 Days AgoSaved
Easy Apply
Remote or Hybrid
7 Locations
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Maintain and improve multi-cloud Kubernetes infrastructure, CI/CD (Argo Workflows/ArgoCD), observability, and networking. Build reliable continuous deployment tooling and onboarding flows, provide internal support, collaborate across Platform Engineering, contribute upstream (open-source/operators), and participate in a 24/7 on-call rotation to resolve deployment infrastructure issues.
Top Skills: AlertingArgo WorkflowsArgocdAWSAzureCi/CdContainersDnsGCPGoKubernetesLinuxLoad BalancerObservabilityPythonService MeshTcp/IpTls
28 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
195K-258K Annually
Senior level
195K-258K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills: Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
Reposted 28 Days AgoSaved
Remote
USA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
An Hour AgoSaved
Remote
United States
Senior level
Senior level
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills: AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
2 Hours AgoSaved
Remote
Mỹ
Entry level
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure, deployment, observability, and reliability capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, Kubernetes, CI/CD, incident response, SLOs, cost optimization, production readiness, automation, and support for hybrid, edge, on-premises, and customer-controlled environments. The role is client-facing, remote, and may involve travel.
Top Skills: AzureAzure ArcAzure DevopsBicepGithub ActionsKubernetesTerraform
2 Hours AgoSaved
Remote
United States
135K-175K Annually
Entry level
135K-175K Annually
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills: AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Reposted YesterdaySaved
Remote
North Carolina, USA
139K-282K Annually
Senior level
139K-282K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Support and optimize large-scale compute, virtualization, and multi-cloud infrastructure. Administer and automate monitoring platform (Splunk, ServiceNow), write SPL queries and dashboards, automate with Python/Ansible/Terraform, manage Linux servers, design and review HLD/LLD and implementation plans, and resolve complex outages to maximize availability and compliance.
Top Skills: AnsibleAWSCentos)Cisco HyperflexCisco UcsItilKubernetesLinux (AlmaPythonServicenowSplSplunkTerraformVMware
YesterdaySaved
In-Office or Remote
2 Locations
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills: AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account