Top Remote Site Reliability Engineer Jobs

2 Days AgoSaved
Remote
United States
145K-200K Annually
Senior level
145K-200K Annually
Senior level
Software
Lead SRE to define strategy and roadmap for reliability, scalability, observability, and automation across cloud and hybrid environments. Design and operate containerized production workloads, infrastructure-as-code, monitoring/alerting, incident management, and compliance for regulated domains. Mentor SREs, partner with security and product teams, manage cloud costs and capacity, and build a developer platform to improve delivery and production stability.
Top Skills: AWSAws MarketplaceAzureAzure MarketplaceGCPGoogle Cloud MarketplaceGrafanaKubernetesPrometheusTerraform
2 Days AgoSaved
Remote
North Carolina, USA
129K-256K Annually
Mid level
129K-256K Annually
Mid level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Manage and support large-scale compute, virtualization, and multi-cloud infrastructure; operate and optimize a monitoring platform (Splunk); automate and script infrastructure tasks (Python, Ansible, Terraform); perform sys-admin duties, patching, security configuration, and compliance reporting; produce design and implementation artifacts and lead incident response and remediation.
Top Skills: AnsibleCentos)Linux (AlmalinuxPythonServicenowSplSplunkTerraform
Reposted 2 Days AgoSaved
Remote
North Carolina, USA
139K-282K Annually
Senior level
139K-282K Annually
Senior level
Cloud • Information Technology • Internet of Things • Professional Services • Software
Support and optimize large-scale compute, virtualization, and multi-cloud environments. Administer and automate monitoring platforms (Splunk, ServiceNow), write SPL, perform Linux sysadmin tasks, automate with Python/Ansible/Terraform, design and review HLD/LLD and change plans, troubleshoot outages, and provide security/compliance reporting.
Top Skills: AlmalinuxAnsibleAWSCentosCisco UcsHyperflexKubernetesMicrosoftPythonServicenowSplunkSplunk SplTerraformVMware
Reposted 2 Days AgoSaved
Remote
United States
147K-168K Annually
Senior level
147K-168K Annually
Senior level
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills: AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Reposted 2 Days AgoSaved
Remote
USA
235K-275K Annually
Expert/Leader
235K-275K Annually
Expert/Leader
Legal Tech • Software
Senior technical leader for SRE driving observability, platform infrastructure, SLIs/SLOs, incident response, automation, and self-service platform capabilities. Shapes reliability strategy, mentors engineers, and ensures production-scale operational excellence.
Top Skills: AiopsBashDatadogGoInfrastructure As CodeKubernetesNew RelicObservabilityPython
2 Days AgoSaved
Remote
30 Locations
Mid level
Mid level
Information Technology
Maintain and optimize mission-critical bare-metal infrastructure, build automation for web2/web3 deployments, perform R&D and monitoring for blockchain validator/RPC/operator nodes, ensure SLIs/SLOs, and participate in on-call rotation.
Top Skills: AnsibleBare-MetalDockerGCPGithub ActionsGoHaproxyKubernetesOracle CloudPostgresPythonTerraform
Reposted 2 Days AgoSaved
In-Office or Remote
17 Locations
Mid level
Mid level
Information Technology • Software • Web3 • Infrastructure as a Service (IaaS)
Operate and improve the Pod platform: respond to incidents, investigate root causes, build automation and observability, design monitoring/alerting, reduce alert fatigue, and drive reliability improvements across production systems.
Top Skills: BashCi/CdCloudDockerGrafanaLinuxPagerdutyPrometheusPythonRust
Reposted 8 Days AgoSaved
In-Office or Remote
New York, NY, USA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
As a Senior Site Reliability Engineer, you will enhance platform reliability, lead incident management, and drive AI-driven improvements in operational workflows.
Top Skills: Amazon Web ServicesDatadogDynamoDBEnvoyEvent Driven ArchitecturesGrpcHTTPIstioJSONKotlinKubernetesLaunchdarklyModern JavaMySQLProtocol BuffersTerraformVitess
Reposted 8 Days AgoSaved
In-Office or Remote
8 Locations
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 4 Days AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills: CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Reposted 4 Days AgoSaved
Remote
USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills: AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
Reposted 4 Days AgoSaved
Remote
USA
110K-140K Annually
Senior level
110K-140K Annually
Senior level
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills: AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 4 Days AgoSaved
In-Office or Remote
2 Locations
80K-133K Annually
Mid level
80K-133K Annually
Mid level
Consulting
Maintain and improve reliability of cloud-based enterprise systems by implementing SRE practices. Participate in design and code reviews, incident management, automation (IaC/CI-CD), monitoring, documentation, and collaboration with cross-functional teams to reduce downtime and improve scalability and security.
Top Skills: Ansible Automation PlatformArtifactoryAWSAzureBashCi/CdGitlabIacLinuxPackerPowershellPythonTerraformWindows
Reposted 4 Days AgoSaved
In-Office or Remote
San Francisco, CA, USA
114K-235K Annually
Mid level
114K-235K Annually
Mid level
Social Media
Operate, scale, and improve a cloud-native platform on AWS and Kubernetes. Manage GitOps deployments with ArgoCD and Helm, provision infra with Terraform/Terragrunt, build CI/CD automation, enhance observability, respond to incidents, reduce operational toil through scripting, and collaborate with security and application teams to improve reliability and platform guardrails.
Top Skills: ArgocdAWSBashContainersEksGithub ActionsGitopsHelmIamKubernetesLinuxPythonTerraformTerragrunt
Reposted 4 Days AgoSaved
Remote
Texas, USA
Mid level
Mid level
Blockchain
The Blockchain Site Reliability Engineer is responsible for maintaining blockchain nodes' reliability, monitoring, incident response, and building automation tools to enhance operations.
Top Skills: DockerElkGoGrafanaJavaScriptKubernetesLinuxPrometheusPythonRustShell
5 Days AgoSaved
Remote
2 Locations
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Top Skills: Argo CdBazelFluxGitopsGoGrafanaInfluxdbKubernetesPrometheusPython
6 Days AgoSaved
In-Office or Remote
4 Locations
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, operate, and improve large-scale Microsoft 365 and Purview services. Automate operational processes, build telemetry and monitoring pipelines, develop scripts/code, troubleshoot and optimize systems, participate in on-call incident response, and collaborate with engineering teams to improve availability, performance, security, and customer experience.
Top Skills: CC#C++JavaJavaScriptMicrosoft 365Microsoft CloudPurviewPython
11 Days AgoSaved
Remote or Hybrid
San Francisco, CA, USA
147K-278K Annually
Senior level
147K-278K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
Reposted 6 Days AgoSaved
Remote
US
101K-161K Annually
Senior level
101K-161K Annually
Senior level
Cloud • Software • Analytics
Join Arista Networks as a Site Reliability Engineer to manage CloudVision service reliability, scalability, and stability in a FedRAMP environment, focusing on areas like architecture, security, and performance optimization.
Top Skills: AnsibleBashGCPGkeGoKubernetesPulumiPython
Reposted 6 Days AgoSaved
In-Office or Remote
5 Locations
Senior level
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Generative AI
The Site Reliability Engineer will develop, deploy, and operate AI infrastructure, focusing on high-performance and scalable machine learning systems using Kubernetes and cloud platforms.
Top Skills: AWSAzureC++GCPGoKubernetesOci
7 Days AgoSaved
Remote
USA
140K-170K Annually
Mid level
140K-170K Annually
Mid level
Artificial Intelligence • Software • Generative AI • Automation
Operate and harden Blitzy's self-hosted, Kubernetes-based AI platform inside customer-controlled secure cloud environments. Own deployments, upgrades, capacity planning, observability, incident response, and customer-facing technical coordination while championing security and feeding operational learnings back into the product roadmap.
Top Skills: Alerting)BashCloud (Aws/Gcp/Azure)Container OrchestrationGoInfrastructure-As-CodeKubernetesMetricsObservability (LoggingPulumiPythonTerraformTracing
7 Days AgoSaved
Remote
2 Locations
Senior level
Senior level
Other
Lead architecture and delivery of highly available, resilient cloud and on‑prem systems for the Password Safe platform. Own platform engineering, CI/CD pipelines, IaC/GitOps, release orchestration, observability (metrics/logs/traces), chaos engineering, SLO/SLI definition, and core services. Mentor engineers, define SRE strategy, and drive reliability, security, and automation improvements.
Top Skills: AnsibleApi GatewaysAWSAzureBlue Green DeploymentsC#CachesCanary DeploymentsChaos EngineeringCi/CdConfiguration ManagementDatadogDevsecopsDockerGitopsGoGrafana CloudJavaKubernetesLinuxOpentelemetryOpentofuSecrets ManagementService MeshTerraformWindows
Reposted 8 Days AgoSaved
In-Office or Remote
3 Locations
192K-455K Annually
Senior level
192K-455K Annually
Senior level
Artificial Intelligence • Information Technology • Cybersecurity • Defense
The Forward Deployed Site Reliability Engineer ensures the reliability of a mission-critical platform, manages incident response, defines SLIs and SLOs, and liaises between engineering and government customers.
Top Skills: AWSBashDockerGrafanaLokiMimirPrometheusPythonTerraform
Reposted 8 Days AgoSaved
In-Office or Remote
3 Locations
Senior level
Senior level
Artificial Intelligence
The Deployment Engineer will build and operate AI inference clusters, ensure scalable deployments, optimize allocation, and maintain infrastructure. Responsibilities include software updates, telemetry development, and collaborative improvements with teams.
Top Skills: DockerGrafanaInfluxdbK8SLinuxPrometheusPython
9 Days AgoSaved
Remote
United States
152K-253K Annually
Senior level
152K-253K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Build and operate the Veeam Data Cloud GOV environment: map systems, write runbooks, define SLIs/SLOs, run incident response, close observability gaps, automate deployments and support fleet management while working across security and compliance constraints.
Top Skills: Api ManagementApplication InsightsArgocdAws CloudformationAzureAzure Arm TemplatesAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account