Top Remote Site Reliability Engineer Jobs

Reposted 11 Days AgoSaved
In-Office or Remote
7 Locations
Mid level
Mid level
Cloud • Software
As a Site Reliability / Gitops Engineer, you will automate operations, develop Infrastructure as Code, maintain core services, and collaborate on service architecture.
Top Skills: Ci/CdCloud ComputingElasticsearchGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 22 Days AgoSaved
Remote
United States
Senior level
Senior level
Logistics • Software
Own and operate scalable infrastructure on GCP (GKE, Cloud Run, AlloyDB); author Terraform modules; manage containerized workloads; build observability in Datadog; design CI/CD in GitHub Actions; automate operational workflows; lead incident response and post-mortems; partner with engineers to improve reliability, cost, and automation.
Top Skills: AlloydbClickhouseCloud RunCloudflare WorkersDatadogDockerGCPGithub ActionsGkeGoGrafanaIamKafkaKubernetesNetworkingPostgresPrometheusPub/SubPythonRedisRedpandaTerraformTypescript
Reposted 11 Days AgoSaved
In-Office or Remote
7 Locations
200K-200K Annually
Senior level
200K-200K Annually
Senior level
Cloud • Software
The Senior Site Reliability / Gitops Engineer will drive automation and collaboration within the IS team, enhancing Canonical's IT operations and services while managing infrastructure as code and cloud technologies.
Top Skills: Cloud ComputingDockerElasticsearchGitopsGrafanaIacKubernetesLinuxPrometheusPython
Reposted 23 Days AgoSaved
Remote
United States
Senior level
Senior level
Insurance
Lead reliability and observability for the financial data platform: define SLOs/SLIs, build metrics pipelines, extend instrumentation across Velocity, Redpanda, MuleSoft, Snowflake, Fabric, and AWS; implement incident management (Datadog -> Incident.io -> ServiceNow), scale automation and remediation, and design AI-assisted SRE agents using Cursor for triage and root-cause analysis.
Top Skills: Ai Coding AssistantsAWSCursorD365DatadogFabricGrafanaIncident.IoKafkaLlmsMulesoftOpentelemetryPower AppsPrometheusRedpandaServicenowSnowflakeVelocity
Reposted 24 Days AgoSaved
Remote
USA
113K-176K Annually
Senior level
113K-176K Annually
Senior level
Other • Social Impact
The Senior Site Reliability Engineer is responsible for maintaining Wikimedia's infrastructure, improving reliability, automating processes, and collaborating with teams. The role involves troubleshooting, managing deployments, and leading incident responses while working remotely.
Top Skills: AnsibleBashCassandraDebianGoGrafanaHhvmKubernetesMariadbMemcachedPHPPrometheusPuppetPythonRedisRubyShell
Reposted 25 Days AgoSaved
Remote
United States
152K-195K Annually
Senior level
152K-195K Annually
Senior level
Information Technology • Security • Cybersecurity
Design, build, and scale Kubernetes-based, multi-tenant infrastructure and CI/CD systems. Own AI tooling infrastructure (MCP servers) and secure AI access patterns. Optimize CI/CD, streaming analytics (Kafka, Flink, ClickHouse), observability, and incident response. Implement IaC (Terraform, Helm, Pulumi), GitOps (Argo CD), progressive delivery, automated testing, and mentor engineering teams.
Top Skills: Ai AgentsAi/Llm ToolingAksArgo CdBashClickhouseDatadogEksFlinkGithub ActionsGitlab CiGitopsGkeGoGrafanaHelmJenkinsKafkaKubernetesLangfuseLangsmithMcp ServersMlopsOpentelemetryPrometheusPulumiPythonTerraform
Reposted 25 Days AgoSaved
In-Office or Remote
Austin, TX, USA
50K-80K Annually
Senior level
50K-80K Annually
Senior level
Artificial Intelligence • Information Technology • Software
The Senior SRE will manage multi-cloud infrastructure, ensuring reliability and scalability. Responsibilities include building CI/CD pipelines, defining SLOs, and implementing automation.
Top Skills: Ai-Assisted DevelopmentAWSAzureClaude CodeDatadogGCPGrafanaKubernetesTerraform
Reposted 26 Days AgoSaved
Remote
US
110K-130K Annually
Senior level
110K-130K Annually
Senior level
Software
Support and improve production SaaS infrastructure across AWS, colocation, and hosted platforms. Administer Windows and Linux systems, virtualization, storage, networking, and database support. Lead incident troubleshooting, root cause analysis, monitoring improvements, vulnerability remediation, automation initiatives, and medium-sized infrastructure projects. Collaborate on compliance (SOX/PCI/HIPAA), disaster recovery, and documentation to increase operational reliability.
Top Skills: AWSBackup And RecoveryBashDatabasesFirewallsLinuxMonitoring PlatformsNetworkingPowershellPythonStorage SystemsVirtualizationWindows Server
27 Days AgoSaved
Remote
US
104K-163K Annually
Senior level
104K-163K Annually
Senior level
Software • Financial Services
Lead reliability, observability, and resilience for cloud-based financial SaaS. Define SLOs/SLIs, design monitoring/tracing, own incident response and runbooks, build IaC and automation, implement AIOps, perform chaos and load testing, and write production-grade Python tooling while ensuring security and compliance.
Top Skills: AnsibleAWSAzureBashCi/CdCloudFormationDatadogElkGitGrafanaNew RelicPowershellPrometheusPythonTerraform
Reposted 27 Days AgoSaved
Remote
United States
Senior level
Senior level
Big Data
You will manage AWS infrastructure, automate deployments, debug application issues, and improve the operational health of Metabase Cloud.
Top Skills: AWSDatadogGoGrafanaKubernetesPrometheusPythonTerraform
28 Days AgoSaved
Remote
USA
134K-184K Annually
Senior level
134K-184K Annually
Senior level
Healthtech
Lead the migration from legacy Azure services to a Kubernetes-based, containerized microservices platform. Design, build, and scale infrastructure, implement observability (monitoring/alerting/logging), drive incident response and SLOs, automate with IaC and CI/CD, optimize cost and networking, mentor teams, and document systems to ensure reliable, scalable healthcare platform operations.
Top Skills: .NetAWSAzureAzure Entra IdBashC#DatadogGCPGithub ActionsGitlab Ci/CdGrafanaHelmKubernetesPrometheusPythonTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
28 Days AgoSaved
Remote
USA
95K-135K Annually
Senior level
95K-135K Annually
Senior level
Real Estate • Financial Services • PropTech
Lead AWS-based SRE activities for products migrated from on-prem: ensure reliability, observability, automation, CI/CD (GitOps), Kubernetes/EKS operations, Terraform IAC, database/RDS administration, networking and security, and collaborate with development and platform teams to optimize SaaS operations.
Top Skills: AmiArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAws Well-Architected FrameworkAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad Balancer (Elb/Alb)PowershellPythonRdsService MeshSQLTerraformWget
28 Days AgoSaved
Remote
13 Locations
Senior level
Senior level
Fintech • Information Technology
Operate and improve brokerage platform reliability: on-call incident response, define SLIs/SLOs, enhance observability, deploy infrastructure via GitOps, and own PostgreSQL performance, migrations, HA/DR, and mentoring.
Top Skills: AlertingDnsGitopsGoKubernetesLinuxLoad Balancing (L4)Load Balancing (L7)LogsMetricsPostgresPythonTlsTracingVpc
One Month AgoSaved
Remote
United States
152K-253K Annually
Senior level
152K-253K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Build and run Gov/Sovereign cloud SRE for Veeam Data Cloud: document platform, define SLIs/SLOs, run incident response, close observability gaps, design resilient Azure infrastructure, automate IaC/CI/CD pipelines, support on-call, and collaborate with security/compliance teams to operationalize reliability.
Top Skills: Application InsightsArgocdAws CloudformationAzureAzure Api ManagementAzure Arm TemplatesAzure DevopsAzure FunctionsAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
One Month AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability and performance efforts for distributed metadata systems: tune and optimize systems, develop monitoring and automation, manage rollouts, troubleshoot incidents, run simulations and analytics, and support database/configuration management to improve global network stability and capacity.
Top Skills: Big DataLinuxPostgresPythonSQLUnix
One Month AgoSaved
In-Office or Remote
2 Locations
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability, scalability, and observability for high-density AI hardware infrastructure. Build Python automation and IaC, design telemetry and Prometheus/Grafana dashboards, implement AI-assisted tooling and anomaly detection, manage 24x7 on-call incident response, and coordinate vendor field operations to ensure uptime.
Top Skills: Bare-MetalBgpGrafanaInfrastructure-As-CodeIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonRest ApisSlackTimeseries Databases
Reposted One Month AgoSaved
Remote
United States
142K-195K Annually
Senior level
142K-195K Annually
Senior level
Software
Design, implement, and operate observability and reliability for cloud platforms. Measure and monitor production systems, reduce toil via automation, drive incident response and on-call practices, and partner with product and platform teams to improve scalability, resiliency, and observability.
Top Skills: AnsibleAWSAzureBlamelessCloud SdksCloudwatchContainersCriblFirehydrantGrafanaJavaScriptKibanaKubernetesLinuxNew RelicNode.jsPagerdutyPrometheusSentrySplunkTerraformTypescript
Reposted One Month AgoSaved
Remote or Hybrid
3 Locations
240K-312K Annually
Senior level
240K-312K Annually
Senior level
Software
Operate and scale Lambda's multi-tenant cloud networking and SDN infrastructure; run Kubernetes control plane and SmartNIC dataplane software; build automation, CI/CD and GitOps workflows; deploy monitoring and observability; collaborate across teams, drive incident response and on-call rotation, capacity planning, and postmortems to improve reliability.
Top Skills: AnsibleCCi/CdDpdkGitopsGoHelmKubernetesLinuxMonitoring/ObservabilityOpenstack NeutronOvnOvsPythonSmartnicsSr-IovTerraform
Reposted One Month AgoSaved
In-Office or Remote
9 Locations
170K-290K Annually
Expert/Leader
170K-290K Annually
Expert/Leader
Artificial Intelligence • Software
As a Software Engineer in Reliability, you'll architect and manage multi-cloud GPU infrastructure, ensuring performance, security, and scale while debugging complex hardware/software issues.
Top Skills: AmdAWSBashGoGpuInfinibandLinuxNvidiaOciPythonRdma
Reposted One Month AgoSaved
Remote or Hybrid
3 Locations
240K-312K Annually
Senior level
240K-312K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Operate and scale a multi-tenant cloud networking platform and SDN infrastructure, manage Kubernetes control plane and SmartNIC dataplane software, build automation and CI/CD/GitOps workflows, deploy observability and monitoring, participate in on-call incident response, and collaborate across software, platform, and networking teams to improve reliability and deployments.
Top Skills: AnsibleCi/CdGitopsKubernetesLinuxPythonSmartnics
Reposted One Month AgoSaved
In-Office or Remote
3 Locations
100K-125K Annually
Senior level
100K-125K Annually
Senior level
Healthtech • Pet • Biotech
Senior SRE responsible for designing and modernizing CI/CD and deployment systems, automating AWS Serverless infrastructure, improving observability and incident response, enforcing release and security practices, and guiding engineering teams to scale resilient global services.
Top Skills: AuroradbAws CloudformationAws LambdaAzure Entra IdCloudfrontDynamoDBEventbridgeGitGitGithub ActionsMavenOauth2Openid ConnectS3SnsSqsTerraform
Reposted 21 Days AgoSaved
Remote or Hybrid
4 Locations
165K-330K Annually
Mid level
165K-330K Annually
Mid level
Software
As a Site Reliability Engineer, you'll build and maintain infrastructure for ML models, automate processes, and collaborate cross-functionally.
Top Skills: Circle CiCloudFormationElk StackGithub ActionsGitlab CiGrafanaJenkinsKubernetesOpentelemetryPrometheusPulumiTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account