Top Remote Site Reliability Engineer Jobs

Reposted One Month AgoSaved
Remote
United States
Senior level
Senior level
Big Data
You will manage AWS infrastructure, automate deployments, debug application issues, and improve the operational health of Metabase Cloud.
Top Skills: AWSDatadogGoGrafanaKubernetesPrometheusPythonTerraform
One Month AgoSaved
Remote
USA
95K-135K Annually
Senior level
95K-135K Annually
Senior level
Real Estate • Financial Services • PropTech
Lead AWS-based SRE activities for products migrated from on-prem: ensure reliability, observability, automation, CI/CD (GitOps), Kubernetes/EKS operations, Terraform IAC, database/RDS administration, networking and security, and collaborate with development and platform teams to optimize SaaS operations.
Top Skills: AmiArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAws Well-Architected FrameworkAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad Balancer (Elb/Alb)PowershellPythonRdsService MeshSQLTerraformWget
One Month AgoSaved
Remote
United States
152K-253K Annually
Senior level
152K-253K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Build and run Gov/Sovereign cloud SRE for Veeam Data Cloud: document platform, define SLIs/SLOs, run incident response, close observability gaps, design resilient Azure infrastructure, automate IaC/CI/CD pipelines, support on-call, and collaborate with security/compliance teams to operationalize reliability.
Top Skills: Application InsightsArgocdAws CloudformationAzureAzure Api ManagementAzure Arm TemplatesAzure DevopsAzure FunctionsAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
One Month AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability and performance efforts for distributed metadata systems: tune and optimize systems, develop monitoring and automation, manage rollouts, troubleshoot incidents, run simulations and analytics, and support database/configuration management to improve global network stability and capacity.
Top Skills: Big DataLinuxPostgresPythonSQLUnix
One Month AgoSaved
In-Office or Remote
2 Locations
121K-219K Annually
Senior level
121K-219K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability, scalability, and observability for high-density AI hardware infrastructure. Build Python automation and IaC, design telemetry and Prometheus/Grafana dashboards, implement AI-assisted tooling and anomaly detection, manage 24x7 on-call incident response, and coordinate vendor field operations to ensure uptime.
Top Skills: Bare-MetalBgpGrafanaInfrastructure-As-CodeIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonRest ApisSlackTimeseries Databases
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
Remote
United States
142K-195K Annually
Senior level
142K-195K Annually
Senior level
Software
Design, implement, and operate observability and reliability for cloud platforms. Measure and monitor production systems, reduce toil via automation, drive incident response and on-call practices, and partner with product and platform teams to improve scalability, resiliency, and observability.
Top Skills: AnsibleAWSAzureBlamelessCloud SdksCloudwatchContainersCriblFirehydrantGrafanaJavaScriptKibanaKubernetesLinuxNew RelicNode.jsPagerdutyPrometheusSentrySplunkTerraformTypescript
Reposted One Month AgoSaved
Remote or Hybrid
3 Locations
240K-312K Annually
Senior level
240K-312K Annually
Senior level
Software
Operate and scale Lambda's multi-tenant cloud networking and SDN infrastructure; run Kubernetes control plane and SmartNIC dataplane software; build automation, CI/CD and GitOps workflows; deploy monitoring and observability; collaborate across teams, drive incident response and on-call rotation, capacity planning, and postmortems to improve reliability.
Top Skills: AnsibleCCi/CdDpdkGitopsGoHelmKubernetesLinuxMonitoring/ObservabilityOpenstack NeutronOvnOvsPythonSmartnicsSr-IovTerraform
Reposted One Month AgoSaved
Remote or Hybrid
3 Locations
240K-312K Annually
Senior level
240K-312K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Operate and scale a multi-tenant cloud networking platform and SDN infrastructure, manage Kubernetes control plane and SmartNIC dataplane software, build automation and CI/CD/GitOps workflows, deploy observability and monitoring, participate in on-call incident response, and collaborate across software, platform, and networking teams to improve reliability and deployments.
Top Skills: AnsibleCi/CdGitopsKubernetesLinuxPythonSmartnics
Reposted One Month AgoSaved
In-Office or Remote
9 Locations
170K-290K Annually
Expert/Leader
170K-290K Annually
Expert/Leader
Artificial Intelligence • Software
As a Software Engineer in Reliability, you'll architect and manage multi-cloud GPU infrastructure, ensuring performance, security, and scale while debugging complex hardware/software issues.
Top Skills: AmdAWSBashGoGpuInfinibandLinuxNvidiaOciPythonRdma
Reposted One Month AgoSaved
In-Office or Remote
3 Locations
100K-125K Annually
Senior level
100K-125K Annually
Senior level
Healthtech • Pet • Biotech
Senior SRE responsible for designing and modernizing CI/CD and deployment systems, automating AWS Serverless infrastructure, improving observability and incident response, enforcing release and security practices, and guiding engineering teams to scale resilient global services.
Top Skills: AuroradbAws CloudformationAws LambdaAzure Entra IdCloudfrontDynamoDBEventbridgeGitGitGithub ActionsMavenOauth2Openid ConnectS3SnsSqsTerraform
Reposted One Month AgoSaved
Remote or Hybrid
5 Locations
165K-330K Annually
Mid level
165K-330K Annually
Mid level
Software
As a Site Reliability Engineer, you'll build and maintain infrastructure for ML models, automate processes, and collaborate cross-functionally.
Top Skills: Circle CiCloudFormationElk StackGithub ActionsGitlab CiGrafanaJenkinsKubernetesOpentelemetryPrometheusPulumiTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account