Top Site Reliability Engineer Jobs

Reposted 9 Days AgoSaved
In-Office
San Jose, CA, USA
149K-361K Annually
Senior level
149K-361K Annually
Senior level
News + Entertainment
As a Senior Machine Learning Engineer, you will develop advanced machine learning and deep learning models and platforms for optimizing advertising performance and conduct complex experiments.
Top Skills: AIControl SystemsDeep LearningMachine LearningReinforcement LearningStatistical Techniques
Reposted 9 Days AgoSaved
In-Office
Austin, TX, USA
Senior level
Senior level
News + Entertainment
Design, operate, and scale cloud-native ML infrastructure across GCP and AWS (GPU/TPU), build CI/CD for models, maintain low-latency real-time inference systems, define observability and monitoring for ML models, participate in on-call incident response, and partner with data scientists to improve MLOps and platform usability.
Top Skills: AerospikeApache AirflowApache FlinkSparkAWSChrononDatadogEksGCPGitlab RunnerGkeGpuGrafanaJavaJenkinsKafkaKubernetesKv StoreMlflowPrometheusPythonRayScalaTerraformTpuVector Database
Reposted 9 Days AgoSaved
Remote
United States
115K-135K Annually
Mid level
115K-135K Annually
Mid level
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills: ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Reposted 9 Days AgoSaved
Hybrid
Palo Alto, CA, USA
175K-229K Annually
Senior level
175K-229K Annually
Senior level
Other
Design, deploy, and operate a high-performance SaaS manufacturing platform on cloud infrastructure. Ensure scalability, reliability, automation, monitoring, and KPIs. Integrate AI into dev/ops workflows, lead impactful projects, and improve uptime, latency, and efficiency. Work cross-functionally in a fast-paced startup while meeting U.S. government data access citizenship requirements.
Top Skills: Ai ToolsApm ToolsAWSContainerizationKubernetesLinuxLogging ToolsMonitoring ToolsShellTerraform
Reposted 9 Days AgoSaved
In-Office
Redmond, WA, USA
165K-230K Annually
Senior level
165K-230K Annually
Senior level
Aerospace • Other
Design, deploy, and scale on-prem Kubernetes clusters and core infrastructure for Starlink. Build automation, manage databases, monitoring, and distributed storage. Collaborate with engineers to improve service lifecycle, availability, and performance; troubleshoot across the Starlink stack and drive reliability improvements.
Top Skills: AnsibleBashBazelC++GoKubernetesLinuxMakefilesOci ContainersPythonTcp/IpTerraform
Reposted 9 Days AgoSaved
In-Office
Redmond, WA, USA
165K-230K Annually
Senior level
165K-230K Annually
Senior level
Aerospace • Other
Design, deploy, and scale on‑premise compute and core infrastructure for Starlink. Develop automation, manage databases/monitoring/distributed storage, collaborate with software teams, troubleshoot end-to-end, and improve deployment and developer velocity.
Top Skills: AnsibleBashCC++DatabasesDistributed StorageDockerGoHypervisor TechnologiesKubernetesLinuxMonitoringPythonTcp/IpTerraformVirtualization
Reposted 9 Days AgoSaved
In-Office
Palo Alto, CA, USA
165K-190K Annually
Mid level
165K-190K Annually
Mid level
Cybersecurity
Ensure reliability, scalability, observability, and cost efficiency of a customer-facing SaaS security platform. Manage Kubernetes/Helm deployments, CI/CD (GitLab/ArgoCD), monitoring, and service verification. Embed with engineering teams, optimize developer CI/CD workflows, monitor and debug production on AWS/GCP, and participate in a 24/7 on-call rotation.
Top Skills: ArgocdAWSGCPGitlab Ci/CdGrafanaHelmKubernetesMicroservicesPrometheus
Reposted 9 Days AgoSaved
In-Office
New York, NY, USA
160K-230K Annually
Mid level
160K-230K Annually
Mid level
Fintech • Payments • Financial Services
Operate and maintain AWS infrastructure and CI/CD pipelines, deploy containerized applications, write Terraform IaC, implement monitoring/observability, troubleshoot incidents, resolve security vulnerabilities, participate in on-call rotation, and produce operational/runbook documentation to ensure resilient cloud platform delivery.
Top Skills: AgileAlbAmazon Web Services (Aws)Aws IamAws LambdaCloudbeaverCloudwatchDevsecopsDirect ConnectDnsEc2EcsEfsEksElbGitlab CiGlueGrafanaHelmJavaJbossJIRANginxOktaOracle RdsPythonRds PostgresRoute53S3ScalaSnsTerraformTomcatTransit GatewayVpc
Reposted 9 Days AgoSaved
In-Office
Mountain View, CA, USA
200K-260K Annually
Senior level
200K-260K Annually
Senior level
Artificial Intelligence • Software • Generative AI
The Lead Site Reliability Engineer will drive technical strategy, ensure high service availability, manage cloud infrastructure, and lead a team to optimize systems and automate processes.
Top Skills: AWSAzureDockerGoogle Cloud PlatformKubernetesTerraform
10 Days AgoSaved
Remote
USA
Entry level
Entry level
Cybersecurity
Assist with monitoring system performance, uptime, and reliability; support incident response, troubleshooting, and root cause analysis; monitor SRE alerts in Slack and report issues; learn cloud platforms; and collaborate with development and DevOps teams.
Top Skills: AWSAzureDevOpsDevsecopsGCPGrafanaKubernetesPrometheusSlack
Reposted 15 Days AgoSaved
Hybrid
4 Locations
149K-282K Annually
Senior level
149K-282K Annually
Senior level
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills: AWSGoKubernetesPuppetPythonTerraform
10 Days AgoSaved
In-Office
Albany, NY, USA
Senior level
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
The Senior Site Reliability Engineer will improve the reliability, scalability, performance, and resilience of enterprise systems. Responsibilities include developing automation and self-healing solutions, reducing operational toil, managing cloud infrastructure, Kubernetes, containers, and microservices, and implementing observability through monitoring, logging, alerting, and distributed tracing. The role also handles incident management, root cause analysis, performance tuning, and capacity planning while collaborating across teams. This is a hybrid position requiring 50% onsite work in Albany or Manhattan, with the first week onsite in Albany.
Top Skills: AWSAzureContainersDynatraceElastic StackElkGCPKubernetesMicroservicesPythonSplunk
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 15 Days AgoSaved
In-Office or Remote
New York, NY, USA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
As a Senior Site Reliability Engineer, you will enhance platform reliability, lead incident management, and drive AI-driven improvements in operational workflows.
Top Skills: Amazon Web ServicesDatadogDynamoDBEnvoyEvent Driven ArchitecturesGrpcHTTPIstioJSONKotlinKubernetesLaunchdarklyModern JavaMySQLProtocol BuffersTerraformVitess
Reposted 15 Days AgoSaved
In-Office or Remote
8 Locations
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 15 Days AgoSaved
Easy Apply
Hybrid
Somerville, MA, USA
Easy Apply
160K-200K Annually
Senior level
160K-200K Annually
Senior level
Enterprise Web • Hardware • Internet of Things • Software
Lead observability and reliability efforts: mentor teams on SLIs/SLOs, maintain triage/remediation workflows, perform incident response, debug production systems, and design core infrastructure and tooling for engineering teams.
Top Skills: AlloyClaude SkillsGemini GemsGoGrafanaKubernetesLokiMimirMongoDBOpentelemetryPostgresPrometheusPromqlTempoTypescript
Reposted 15 Days AgoSaved
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills: .NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Reposted 10 Days AgoSaved
Hybrid
Palo Alto, CA, USA
186K-256K Annually
Senior level
186K-256K Annually
Senior level
Automotive • Cloud • Hardware • Software
Lead and operate ProdOps as senior SRE: coordinate and communicate during high-severity incidents, run blameless post-incident reviews, prioritize and drive systemic action items, verify fixes, build automation and observability (including AI-agent workflows), and grow a lean team while mentoring engineers and setting technical standards.
Top Skills: Ai AgentsCloudDatadogGoInstrumentationLlmsLoggingMetricsObservabilityPythonSliSloTracing
Reposted 10 Days AgoSaved
In-Office
Reston, VA, USA
148K-264K Annually
Mid level
148K-264K Annually
Mid level
Cloud • Fintech • HR Tech
Design, implement, test, deploy, and maintain automation and configuration management for containerized services. Automate deployments and scaling on Kubernetes, build self-service platforms for developers, monitor and triage production incidents, run retrospectives, and participate in infrequent on-call rotations. Collaborate across teams, use observability tools to debug, mentor engineers, and drive continuous improvement and operational efficiency.
Top Skills: AnsibleArgocdAWSBashGCPGoGrafanaJenkinsKubernetesPrometheusPythonRubySplunk
11 Days AgoSaved
In-Office
3 Locations
170K-225K Annually
Senior level
170K-225K Annually
Senior level
Fintech • Software • Financial Services
Own AWS and Kubernetes reliability, deployments, observability, incident response, infrastructure as code, CI/CD, and compliance controls for SOC 2 and PCI-DSS. Build scalable platform foundations, AI infrastructure, and greenfield product infrastructure while improving cost efficiency, security, vendor flexibility, and deployment consistency. Participate in on-call support and partner with security and engineering teams across backend, mobile, data, ML, and AI.
Top Skills: Ai/Llm ToolingAmazon EksAWSAws CdkCi/CdDatadogKotlinKubernetesMcpNode.jsOpentelemetryOpentofuPci-DssPythonSoc 2TerraformTypescript
11 Days AgoSaved
In-Office or Remote
2 Locations
138K-171K Annually
Junior
138K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale cloud infrastructure.
Top Skills: AnsibleAWSBashBitbucketDatabricksDatadogDigicertGradleJenkinsNode.jsPythonSplunkTenableTerraformThreatmetrix
11 Days AgoSaved
In-Office
2 Locations
Senior level
Senior level
Artificial Intelligence • Information Technology
Lead the reliability, stability, observability, and scalability of Seekr’s software platform. Design and implement systems, networks, and services that meet SLAs; develop monitoring, automation, and failure-detection solutions; conduct load testing; manage observability tools; troubleshoot production and deployment issues; and support incident response. Collaborate closely with software engineering teams across hybrid cloud and on-premises environments.
Top Skills: AerospikeAnsibleArgocdBashChefDockerElasticsearchElkGitGitlabGrafanaInfluxdbJavaKafkaKubernetesLinuxPrometheusPuppetPythonRubyTerraform
Reposted 11 Days AgoSaved
Remote or Hybrid
United States
150K-225K Annually
Senior level
150K-225K Annually
Senior level
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills: AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
11 Days AgoSaved
In-Office
3 Locations
125K-168K Annually
Senior level
125K-168K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services • Data Privacy
Leads reliability and operational excellence for an enterprise Kubernetes container platform. Responsibilities include incident response, root cause analysis, L3 support, upgrades, observability, automation, capacity planning, performance tuning, security controls, compliance, resilience testing, and runbook development. The role partners with engineering, architecture, security, infrastructure, product, and operations teams while mentoring SRE resources and improving platform adoption, developer experience, and production reliability.
Top Skills: AnsibleArgocdBashBitbucketCi/CdDnsDynatraceElkGitGitlabGitopsGoGrafanaHelmInfrastructure As CodeJenkinsKubernetesLinuxLoad BalancingNetwork PoliciesOpenshiftOpenshift VirtualizationOpentelemetryPrometheusPythonRancherRbacService MeshSplunkTerraformVcfVksVMware
11 Days AgoSaved
In-Office or Remote
7 Locations
Mid level
Mid level
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
Top Skills: AkamaiAmqpAnsibleAWSBashCdnCloudflareDnsDockerGCPGitGitlab CiGrafanaHttp/HttpsKubernetesLinuxPodmanPrometheusPythonRabbitMQTerraformVictoriametricsWafZabbix
Reposted 11 Days AgoSaved
In-Office
Golden, CO, USA
103K-136K Annually
Senior level
103K-136K Annually
Senior level
Manufacturing
The Site Reliability Engineer will ensure the reliability, security, and support of Databricks applications while collaborating with various teams to optimize data workflows and incident management.
Top Skills: AzureCi/CdDatabricksDelta LakePysparkPythonSQLUnity Catalog
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account