Top Site Reliability Engineer Jobs

Reposted 19 Days AgoSaved
In-Office
2 Locations
157K-239K Annually
Senior level
157K-239K Annually
Senior level
Aerospace • Defense
Own and operate Loft's cloud and hybrid network infrastructure: design architectures, secure connectivity, manage network-as-code (IaC), define SLOs and observability, automate toil, and participate in SRE incident response and platform reliability.
Top Skills: ArgocdCi/CdDnsDockerFirewallingFluxcdGCPGitGitopsGrafanaHybrid ConnectivityK8SKubernetesPeeringSdnSite-To-Site RoutingSoftware-Defined NetworkingTerraformVpcVpn
Reposted 19 Days AgoSaved
In-Office
Aliso Viejo, CA, USA
146K-219K Annually
Senior level
146K-219K Annually
Senior level
Gaming
The role involves ensuring production quality, owning system reliability, and participating in decision-making. Responsibilities include incident response and lifecycle management in cloud gaming technologies.
Top Skills: BashC++ElasticsearchGoIstioJavaKafkaKong Api GatewayKubernetesKumaLinkerdMongoDBMySQLPostgresPythonRedisRust
Reposted 19 Days AgoSaved
Remote
United States
205K-270K Annually
Senior level
205K-270K Annually
Senior level
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills: ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
20 Days AgoSaved
Remote
USA
Senior level
Senior level
Information Technology • Professional Services
Provide senior SRE expertise to improve reliability, scalability, performance, and resilience of a cloud-hosted geospatial platform. Design monitoring/observability, automate deployments, support incident response, optimize capacity and performance, and collaborate across DevSecOps, Kubernetes, database, and support teams in a mission-focused DoD environment.
Top Skills: Alerting ToolsArcgis EnterpriseAw S Cloud OneAWSCi/CdContainerized SystemsEsriInfrastructure-As-CodeKubernetesLinuxLogging ToolsMonitoring ToolsRmfScriptingStig
Reposted 25 Days AgoSaved
Remote or Hybrid
2 Locations
110K-155K Annually
Senior level
110K-155K Annually
Senior level
Information Technology • Insurance • Software
Own and operate production services end-to-end to ensure reliability, scalability, performance, and operational health. Define SLIs/SLOs, perform incident response and root cause analysis, build automation and self-healing, manage production changes, and collaborate with engineering, product, and operations teams to improve system design and observability.
Top Skills: .NetAWSC#Ci/CdInfrastructure As CodeJavaKubernetesLinuxPythonReactRelational DatabasesWindows
Reposted 20 Days AgoSaved
In-Office
4 Locations
208K-269K Annually
Senior level
208K-269K Annually
Senior level
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills: Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Reposted 20 Days AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills: AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
Reposted 20 Days AgoSaved
In-Office
Washington, DC, USA
174K-239K Annually
Senior level
174K-239K Annually
Senior level
Cloud
Design, build, and maintain secure, air-gapped cloud platform services and CI/CD pipelines for Okta Federal. Operate mission-critical infrastructure, monitor SLOs/SLIs, run incident response and POA&M remediation, and support Authority to Operate activities. Advocate SRE/DevOps practices across teams while working autonomously in secure facilities.
Top Skills: Air-Gapped EnvironmentsAmazon CloudwatchAws Transit GatewayAws VpcBgpCi/CdEcs FargateEksGrafanaIpsecPythonSplunkTerraformVpc Endpoints
21 Days AgoSaved
In-Office
Austin, TX, USA
Senior level
Senior level
Blockchain • Fintech • Cryptocurrency
Lead design, implementation, and operation of reliable hybrid cloud and on-prem infrastructure. Deliver complex initiatives, build immutable infrastructure with IaC, improve observability and incident response, mentor engineers, and ensure reliability, scalability, and compliance in change-controlled environments.
Top Skills: AnsibleAWSDatadogGithub ActionsHelmKubernetesKustomizeMakefilesPythonTerraformVirtualization
Reposted 21 Days AgoSaved
Hybrid
Rockville, MD, USA
112K-150K Annually
Mid level
112K-150K Annually
Mid level
Artificial Intelligence • Cloud • Software • Cybersecurity
Operate and tune AWS environments to meet SLAs, build observability and alerts, automate infrastructure with IaC and CI/CD, define SLIs/SLOs, support security/compliance within a FISMA Moderate boundary, design resilience and DR plans, and own incident response and post-mortems.
Top Skills: AnsibleAWSAws CloudwatchAws Trusted AdvisorCi/CdCloudFormationDockerGitlab CiJenkinsNew RelicPythonSplunkTerraform
Reposted 21 Days AgoSaved
Hybrid
Atlanta, GA, USA
Mid level
Mid level
Fintech • Payments • Financial Services
Build, operate, and scale AWS-based infrastructure using IaC (Terraform), manage EKS and serverless environments, create CI/CD pipelines, implement observability (OpenTelemetry/Prometheus/New Relic), support Postgres/RDS (Aurora), lead incident response and define SRE practices (SLIs/SLOs/error budgets).
Top Skills: AuroraAWSAws RdsAzureCloudFormationEcsEksGithub ActionsGitlabGoGCPJavaKubernetesNew RelicOpentelemetryOpentofuPostgresPrometheusPythonRubyServerlessTerraformTerragrunt
Reposted 21 Days AgoSaved
In-Office
St. Louis, MO, USA
100K-120K Annually
Senior level
100K-120K Annually
Senior level
Fintech • Analytics
As a Senior Site Reliability Engineer, you'll lead incident recovery, enhance production stability, automate processes, and collaborate with development teams to improve operational efficiency.
Top Skills: AWSAzureBigpandaCloud-Native ApplicationsDatadogDnsDockerGitHTTPKubernetesShell ScriptingTcp/IpUnix
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 21 Days AgoSaved
In-Office
San Francisco, CA, USA
Mid level
Mid level
Enterprise Web • Information Technology • Software
As a Platform Engineer, you will enhance reliability and performance, design operational processes, and build monitoring systems while collaborating with a talented team.
Top Skills: AIAssistantsBackendDeveloper ToolsFrontendInfrastructureMcpsMonitoring SystemsSkills
Reposted 21 Days AgoSaved
In-Office
San Francisco, CA, USA
Mid level
Mid level
Software
Join a passionate team to enhance reliability and performance of the AI control plane, manage deployments, and respond to production incidents while ensuring service quality for customers.
Top Skills: Ai Control PlaneDeveloper ToolsInfrastructure
Reposted 21 Days AgoSaved
Remote
Illinois, USA
150K-224K Annually
Senior level
150K-224K Annually
Senior level
Legal Tech • Software
Lead Site Reliability Engineer responsible for platform availability and reliability of RelativityOne. Drive SRE best practices, build tools, lead projects, coach SREs, work with stakeholders, support incidents, run postmortems, and improve monitoring, automation, and operational efficiency.
Top Skills: Ci/CdDevOpsJenkinsJIRAKubernetesAzureMonitoring And AlertingNew RelicNoSQLPowershellRelativity ServerRelativityoneSQLTableau
Reposted 21 Days AgoSaved
In-Office
Birmingham, AL, USA
Mid level
Mid level
Automotive • Hardware • Logistics
Builds and supports large-scale, distributed, fault-tolerant systems to improve reliability and automation. Administers networks and databases, monitors system health, configures load and data communications, coordinates equipment and vendor orders, and participates in change management to reduce incidents and support cloud transformations.
Top Skills: CloudDistributed ComputingInternet SecurityMonitoring ToolsOracle ErpUnixVersion Control SystemsWindows 2000Windows 98Windows Nt
Reposted 21 Days AgoSaved
Hybrid
3 Locations
220K-235K Annually
Senior level
220K-235K Annually
Senior level
Information Technology • Consulting
As a Senior Staff Site Reliability Engineer, you will lead the SRE team, advocate best practices, ensure resilience in cloud architecture, and mentor team members.
Top Skills: ArgocdCircleCIGoogle Cloud PlatformKubernetesPulumiTerraformTypescript
Reposted 21 Days AgoSaved
In-Office
Miami, FL, USA
Senior level
Senior level
Healthtech
The Senior Software Engineer will enhance system reliability, manage Kubernetes and AWS environments, oversee incident responses, and implement observability measures.
Top Skills: AWSCloudwatchElbGithub ActionsKubernetesObservability ToolingTerraformVpc
Reposted 22 Days AgoSaved
In-Office
Alpharetta, GA, USA
Senior level
Senior level
Fintech • Financial Services
As a Site Reliability Engineer for AI platform, you will ensure system reliability, develop automation, manage infrastructure, and evaluate new technologies.
Top Skills: AnsibleApache KafkaAWSAzureCloudFormationDatadogDockerEfkElkGoGCPGpu ClustersGrafanaHelmJavaKubernetesPrometheusPythonSnowflakeSparkTerraform
22 Days AgoSaved
In-Office or Remote
San Mateo, CA, USA
240K-300K Annually
Senior level
240K-300K Annually
Senior level
Artificial Intelligence • Hardware • Robotics • Software
Build, operate, and scale production cloud infrastructure and Kubernetes/EKS clusters on AWS. Own Terraform-defined infra, CI/CD, observability, networking, and reliability. Troubleshoot production incidents, participate in on-call rotations, automate operational work with Python/Go, and expand multi-region deployments.
Top Skills: Argo CdAWSDatadogEksGithub ActionsGitlab Ci/CdGitopsGoHelmIamJenkinsKubernetesLinuxPostgresPythonSpinnakerTerraformVpc
22 Days AgoSaved
In-Office
San Jose, CA, USA
105K-155K Hourly
Junior
105K-155K Hourly
Junior
Software
Front-line SRE covering 8AM–8PM PST monitoring GPU clusters, networking, storage, and sensors. Execute runbooks, perform hardware triage and physical DC tasks (rack, cable, swap), collect diagnostics for escalation, manage tickets (ServiceNow/Jira), update runbooks, and tag incidents to train the AIOps platform.
Top Skills: AiopsBmcDcgmGpuGrafanaJira Service ManagementLinuxNagiosPrometheusServicenow
22 Days AgoSaved
Remote
United States
180K-240K Annually
Senior level
180K-240K Annually
Senior level
Artificial Intelligence • Marketing Tech • Mobile • Software
Lead design and implementation of scalable, reliable platform systems; define SLIs/SLOs and observability; drive cross-team strategic initiatives; mentor engineers; own production standards, incident management, and cost/operational optimization to improve platform reliability and scalability.
Top Skills: GoJavaPythonTypescript
Reposted 22 Days AgoSaved
In-Office
2 Locations
140K-170K Annually
Senior level
140K-170K Annually
Senior level
Financial Services
Design, build, and operate reliable cloud infrastructure and networking (multi-account AWS, VPC, IAM). Implement IaC, CI/CD pipelines, observability (logging/metrics/alerting), automation, and reliability guardrails. Provide production support and incident response, perform root cause analysis, and collaborate with application teams to co-own system design and continuous improvement, using AI-assisted tools where appropriate.
Top Skills: .NetAi-Assisted Tools (Claude CodeAWSAws OrganizationsBashCi/CdCloudFormationElastic StackGitGithub CopilotIamInfrastructure As CodeJavaJenkinsNode.jsObservabilityOpensearchPowershellPythonTerraformVpcWindsurf)
Reposted 27 Days AgoSaved
Hybrid
San Francisco, CA, USA
160K-250K Annually
Senior level
160K-250K Annually
Senior level
Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Lead design and delivery of scalable cloud infrastructure for the Spend product. Embed with development teams to drive reliability, performance, observability, incident response, and automation. Own SLOs, runbooks, DevOps metrics, and collaborate with central DevOps and security teams to ensure compliance and resilience. Lead infrastructure projects including new service launches, data centre migrations, and modernising data pipelines.
Top Skills: Analytics PipelinesAWSData StreamingDevOpsGCPIncident ResponseKubernetesObservabilitySlosSre
22 Days AgoSaved
Remote
United States
125K-150K Annually
Senior level
125K-150K Annually
Senior level
Information Technology
Design, build, and operate a reliable, scalable developer platform and cloud infrastructure (GCP/GKE). Lead SRE practices: SLO/SLI, observability, incident response, on-call, automation, security-by-default, Terraform IaC, CI/CD, capacity planning, mentoring, and platform enablement across teams.
Top Skills: Ci/CdConfluenceDatadogDockerGCPGitGithub ActionsGkeGoGrafanaGsm (Gcp Secrets Manager)HclHelmHoneycombJIRAKubernetesNew RelicNode.jsOpentelemetryPrometheusPythonSslTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account