Top Site Reliability Engineer Jobs

Reposted 2 Months AgoSaved
Remote
USA
Senior level
Senior level
Information Technology • Cryptocurrency
The Site Reliability Engineer will lead technical initiatives, architect solutions, troubleshoot issues, mentor team members, and improve observability practices.
Top Skills: ArgocdBashElk StackGCPGoGrafanaHelmKubernetesPrometheusPythonTerraform
Reposted 2 Months AgoSaved
Remote
United States
154K-231K Annually
Senior level
154K-231K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Lead technical design and architecture of large-scale Ceph clusters, drive capacity planning and major upgrades, resolve complex cross-functional production incidents, establish automation and reliability standards, and mentor engineers on advanced Ceph operations to scale the global storage platform.
Top Skills: AnsibleBluestoreCephCephfsCinderCrushCsiDdnErasure CodingGoIsilonKubernetesManilaMdsMgrMonNetappNeutronNovaOpenstackOsdPlacement GroupsPurePythonRbdRbd MirroringRgwRgw MultisiteRookS3SaltstackSwiftTerraformVastWeka
2 Months AgoSaved
In-Office
Los Angeles, CA, USA
100K-200K Annually
Mid level
100K-200K Annually
Mid level
Energy • Chemical • Utilities • Manufacturing
Design, implement, and maintain observability, alerting, and developer productivity systems for production and internal services. Instrument services with metrics, logs, and traces, run on-call, respond to incidents, lead reviews, and automate operational workflows to improve reliability and reduce toil.
Top Skills: Ci/CdDatadogDnsGrafanaHTTPInfrastructure-As-CodeLoggingMetricsOpentelemetryPrometheusTlsTracing
Reposted One Month AgoSaved
Hybrid
3 Locations
145K-200K Annually
Senior level
145K-200K Annually
Senior level
Blockchain • Energy • Cryptocurrency
Hands-on role to assess, implement, test, and document backup, restore, failover, and recovery capabilities. Inventory critical systems, design and automate backup and restoration, run recovery exercises, produce runbooks, validate recoverability, measure RTO/RPO, and train system owners. Collaborate with Security, SRE, DevOps, QA, and application teams to harden shared recovery capabilities and transfer operational ownership.
Reposted One Month AgoSaved
Hybrid
Bellevue, WA, USA
120K-150K Annually
Senior level
120K-150K Annually
Senior level
Healthtech • Software • Analytics • Business Intelligence
Senior SRE responsible for designing, building, and operating reliable, scalable distributed systems; owning production reliability (SLOs/SLIs, incident response, MTTR reduction); automating toil with software and platform tooling; driving observability, capacity planning, and cross-team reliability improvements; mentoring engineers and running blameless postmortems.
Top Skills: AWSAzureDockerGCPGithub ActionsGoGrafanaJavaKubernetesOpentelemetryPrometheusPythonTerraformTypescript
Reposted One Month AgoSaved
In-Office
San Francisco, CA, USA
210K-240K Annually
Senior level
210K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Software • Big Data Analytics
The Senior Site Reliability Engineer will design and maintain scalable infrastructure, improve system reliability, manage CI/CD pipelines, and collaborate across teams for operational excellence.
Top Skills: AnsibleArgocdAWSBashDatadogDockerElkGithub ActionsGrafanaKubernetesLinuxOpentelemetryPrometheusPythonTerraform
2 Months AgoSaved
In-Office or Remote
3 Locations
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Embed with customer teams running large-scale GPU training and inference to onboard, tune, debug, and improve reliability. Diagnose fabric, driver, scheduler, and application failures; profile performance; build automation, monitoring, and preflight checks; lead incident response; and convert field learnings into product improvements and reusable reference configurations.
Top Skills: AnsibleBashCgroupsContainer RuntimesCudaDcgmDevice PluginsFabric ManagerGoGpfsHelmInfinibandKubernetesKv CacheLustreNamespacesNcclNvidia DriversNvidia-SmiNvlinkPythonRoceSlurmSshTerraformTopology-Aware SchedulingVastWeka
2 Months AgoSaved
In-Office or Remote
2 Locations
174K-238K Annually
Senior level
174K-238K Annually
Senior level
Cloud
Lead design, automation, and operation of large-scale cloud services for federal customers. Improve reliability via SRE practices (SLIs/SLOs, incident response), build automation and self-service platforms using Go/Python, Terraform, and Kubernetes, and ensure compliance with FedRAMP/IL6. Mentor teams, drive reliability initiatives, and modernize deployments with CI/CD and GitOps.
Top Skills: ArgocdAWSCassandraCi/CdDnsGCPGitopsGoHelmIamKubernetesLoad BalancingMySQLOpensearchPostgresPythonRedisSecrets ManagementTerraformTls
2 Months AgoSaved
In-Office or Remote
Jupiter, FL, USA
Mid level
Mid level
Cloud • Other
Maintain and monitor Rocket.net hosting platform reliability and performance. Provide advanced escalation-level technical support for WordPress VIP customers, troubleshoot Linux-based production environments, web servers, databases, caching, DNS, and networking. Participate in incident response, root cause analysis, automation, documentation, and collaborate with support and engineering teams to improve platform stability.
Top Skills: ApacheBashCachingCdnCloudflareDatadogDnfDnsHTTPHttpsLinuxMariadbMySQLNedataNginxPhp-FpmRedisSshSsl/TlsWafWordpressYum
2 Months AgoSaved
Remote
USA
75K-90K Annually
Mid level
75K-90K Annually
Mid level
3PL: Third Party Logistics
Own and improve uptime for backend services, APIs, workers, and ML pipelines; build monitoring, alerting, auto-remediation, optimize GCP infrastructure and Postgres performance, run on-call rotation and blameless postmortems.
Top Skills: Apollo ServerBashCi/CdCloud RunCloud SqlDatadogDockerExpressGCPGcsGrafanaNext.JsNode.jsPostgresPrometheusPythonReactTerraformTypescriptZabbix
Reposted 2 Months AgoSaved
In-Office
Wacker, IL, USA
132K-220K Annually
Expert/Leader
132K-220K Annually
Expert/Leader
Financial Services
The Staff Site Reliability Engineer will lead Platform Engineering's SRE efforts by defining technical strategy, overseeing architecture, and enhancing operational excellence through mentorship and governance.
Top Skills: ArgocdGCPGkeGoKafkaNode.jsPythonTerraform
Reposted One Month AgoSaved
In-Office
Palo Alto, CA, USA
165K-280K Annually
Senior level
165K-280K Annually
Senior level
Aerospace • Other
Design, deploy, and operate highly available, sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve performance, and advance deployment, monitoring, and alerting. Collaborate across teams through full software lifecycle to build scalable, operable services for Starlink.
Top Skills: AlertingApache FlinkApache KafkaSparkBare Metal Compute ClustersC#Continuous IntegrationGoHbaseHdfsIstioJavaKubernetesLinuxMonitoringPythonScalaVersion Control
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
In-Office
Hawthorne, CA, USA
165K-265K Annually
Senior level
165K-265K Annually
Senior level
Aerospace • Other
Design, upgrade, and operate large-scale distributed systems for Starlink. Improve sharding, geo-redundancy, multi-region deployment, monitoring, and performance. Manage petabyte-scale bare-metal clusters and collaborate across teams through the full software lifecycle.
Top Skills: Apache KafkaSparkC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Reposted One Month AgoSaved
In-Office
Redmond, WA, USA
165K-270K Annually
Senior level
165K-270K Annually
Senior level
Aerospace • Other
Design, deploy, and operate sharded, geo-redundant distributed systems and multi-region infrastructure. Manage petabyte-scale bare-metal clusters, improve deployment/monitoring/alerting, collaborate across teams, and optimize performance throughout the software lifecycle.
Top Skills: Apache KafkaSparkBare MetalC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Reposted 2 Months AgoSaved
In-Office
El Segundo, CA, USA
170K-195K Annually
Mid level
170K-195K Annually
Mid level
Information Technology
Owner of production reliability across cloud and edge: define and drive SLIs/SLOs, build observability (Grafana/Prometheus/Loki/OpenTelemetry), participate in on-call and incident response, encode reliability in infrastructure-as-code (Terraform/OpenTofu), manage Kubernetes clusters, AWS hardening, HA databases, and oversee IoT/edge device fleet operations.
Top Skills: AWSGitGrafanaIamIotKubernetesLokiNebulaNvidia JetsonOpentelemetryOpentofuPrometheusPyrraSlothStatefulsetsTailscaleTerraformWireguard
Reposted 2 Months AgoSaved
In-Office
Alpharetta, GA, USA
Senior level
Senior level
Information Technology • Consulting
Design, build, customize, and support Oracle E-Business Suite and related financial applications. Gather requirements, implement EBS customizations and integrations, optimize performance and SQL, perform testing and upgrades, and provide tier-3 support and issue resolution with Oracle Support.
Top Skills: Invoice AutomationOracle CloudOracle E-Business SuiteOracle Ebs ApisOracle FinancialsOracle Integration CloudOracle Supplier Portal CloudSQLWebcenter Content Imaging
Reposted 2 Months AgoSaved
In-Office
Cambridge, MA, USA
160K-205K Annually
Senior level
160K-205K Annually
Senior level
Software
Design, build, and operate multi-account cloud infrastructure using IaC. Automate customer deployments, manage CI/CD, troubleshoot production across infra/data/app layers, and handle networking, security, and compliance for regulated environments while collaborating with platform and professional services teams.
Top Skills: AirflowAuth0AWSAzureDbtDockerEcsGCPGithub ActionsLlmsOktaPackerPostgresSnowflakeTailscaleTerraformWireguard
Reposted One Month AgoSaved
In-Office
Seattle, WA, USA
Senior level
Senior level
Other
The Sr. Site Reliability Engineer will maintain and administer enterprise systems, troubleshoot operational issues, and develop scripts. This role requires collaboration across teams and participation in project planning and execution.
Top Skills: AnsibleApacheAzureC#ChefIisJavaJbossPerlPowershellPuppetPythonRubyTomcat
Reposted One Month AgoSaved
In-Office
Washington, DC, USA
Senior level
Senior level
Big Data • Analytics • Business Intelligence • Big Data Analytics
Seeking a Site Reliability Engineer to manage AI platform reliability, automate tasks, optimize ML pipelines, and lead incident response in a hybrid engineering role.
Top Skills: ArgocdBigQueryCloud BuildDockerDvcGithub ActionsGoGrafanaKubeflowKubernetesMlflowPrometheusPub/SubPythonTerraformVertex Ai
Reposted One Month AgoSaved
In-Office or Remote
8 Locations
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills: AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
196K-235K Annually
Senior level
196K-235K Annually
Senior level
Artificial Intelligence • Big Data • Software
Own and improve infrastructure for the Data Replication platform: Kubernetes, CI/CD, secrets, networking, cloud (AWS/GCP). Drive reliability, observability, AI-augmented tooling, canary rollouts, incident reduction, runbooks, and partner with product engineers.
Top Skills: Agentic FrameworksAirbyteAWSCdksCi/CdConnector-Based ArchitecturesDatadogGCPGrafanaHelmJavaKubernetesLlmsPrometheusPythonSecrets ManagementTerraform
Reposted One Month AgoSaved
Remote or Hybrid
Location, WV, USA
Senior level
Senior level
Payments
Senior SRE responsible for ensuring high availability and resiliency of a global payments platform by building observability, automations, AI-driven remediation, incident response, and self-healing workflows; participates in on-call rotation and hybrid Philadelphia-based work.
Top Skills: AiopsAksAnthropic (Claude)ApmAzureAzure Ai (Foundry)Azure Sre AgentCi/CdDatadogDnsDynatraceHTTPHttpsIisKubernetesLoad BalancingNew RelicOpenai (Codex)Pagerduty Process AutomationPowershellPythonRundeckSQLT-SqlTcp/IpVMwareWindows Server
One Month AgoSaved
In-Office
Yamato, Boca Raton, FL, USA
105K-175K Annually
Senior level
105K-175K Annually
Senior level
Information Technology • Legal Tech • Analytics
Lead cloud cost optimization and reliability initiatives across Azure and AWS. Build automation, Infrastructure as Code, CI/CD, GitOps, observability, and self-service capabilities using Kubernetes and Terraform. Support AI, machine learning, Generative AI, and GPU-enabled platforms while implementing governance and financial accountability. Mentor engineers, evaluate technologies, lead proofs of concept, and drive platform modernization and operational excellence.
Top Skills: AWSAzureCi/CdFinopsGenerative AiGitopsGpu ComputingInfrastructure As CodeKubernetesMachine LearningObservabilityTerraform
Reposted One Month AgoSaved
Hybrid
San Francisco, CA, USA
Senior level
Senior level
Artificial Intelligence • Healthtech
Own production reliability and performance for a healthcare AI platform: define SLIs/SLOs, run incident response and postmortems, build observability and automation, optimize scalability across serverless and container services, and partner with engineering, security, and compliance to improve operational maturity.
Top Skills: Amazon EcsAurora PostgresAws LambdaBashClickhouseCloudwatchDatadogGithub ActionsGrafanaOpentelemetryPostgresPythonSentrySIEMVanta
Reposted One Month AgoSaved
Remote
2 Locations
150K-195K Annually
Senior level
150K-195K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills: DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account