Top Site Reliability Engineer Jobs

Reposted 23 Days AgoSaved
Remote
USA
103K-287K Annually
Senior level
103K-287K Annually
Senior level
3D Printing • Artificial Intelligence • Software • Design
Lead design and operation of scalable, multi-tenant spatial streaming platforms. Build Terraform-based cloud infrastructure, optimize CDN/content delivery, implement observability (SLI/SLO), run incident response/on-call, conduct post-mortems, enforce compliance and security practices, and mentor DevOps engineers to improve reliability and production readiness.
Top Skills: Aws FargateCdnCoreweaveGrafanaKubernetesPrometheusTerraform
Reposted 23 Days AgoSaved
Remote or Hybrid
4 Locations
165K-330K Annually
Mid level
165K-330K Annually
Mid level
Software
As an AI Support Engineer, you'll manage support requests, resolve user issues, optimize ML models, and contribute to product development.
Top Skills: Tensorrt
Reposted 23 Days AgoSaved
Remote or Hybrid
7 Locations
Senior level
Senior level
Artificial Intelligence • Information Technology • Software
Build and operate the core infrastructure for Arena's online evaluation systems: design low-latency, high-reliability APIs and gateways, implement enterprise-grade features (rate limiting, auth, metering, audit logging), instrument observability (tracing, latency, usage tracking), integrate with LLM providers and the evaluation platform, and collaborate with research and product teams to scale and harden systems for bursty, unpredictable traffic.
Top Skills: Anthropic ApiAWSDistributed TracingGCPGoGoogle Llm ApisKubernetesOpenai ApiPostgresRedisRustTerraform
Reposted 23 Days AgoSaved
In-Office or Remote
2 Locations
165K-225K Annually
Senior level
165K-225K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Build and operate production-grade AI infrastructure using Kubernetes, ensuring high availability, reliability, and performance. Develop custom operators and implement automation for efficient operations and monitoring.
Top Skills: AnsibleBashElk StackEnterprise Storage SystemsGrafanaHigh-Performance NetworkingKubernetesLinuxNvidia Gpu TechnologiesPrometheusPythonTerraform
Reposted 23 Days AgoSaved
In-Office
2 Locations
Senior level
Senior level
Fintech • Analytics
As a Site Reliability Engineer, you will ensure the reliability and performance of a FX trading platform, develop automation, improve system health, and manage SLOs while collaborating with development teams.
Top Skills: AWSAzureBashC#JavaKubernetesPythonSQL
Reposted 23 Days AgoSaved
Remote
USA
Senior level
Senior level
Gaming • Software
The Site Reliability Engineer will manage infrastructure stability and scalability, lead cloud migrations, and optimize performance across systems while mentoring team members.
Top Skills: AnsibleAWSAzureBashChefCloudFormationDatadogDockerElk StackGCPGoGrafanaKubernetesPrometheusPuppetPythonTerraformUnix/Linux
Reposted 23 Days AgoSaved
In-Office
San Francisco, CA, USA
174K-239K Annually
Senior level
174K-239K Annually
Senior level
Cloud
Design, build, and maintain cloud platform services for sensitive federal missions. Operate and monitor air-gapped, secure environments; build CI/CD pipelines without internet; maintain SLOs/SLIs; own runbooks and incident response; support POA&M and Authority to Operate activities; and deliver internal platform enablement while ensuring strict compliance and security.
Top Skills: Aws VpcsBgpCi/CdCloudwatchContainersEcs FargateEksGrafanaIpsecLinuxPythonSplunkTerraformTgwsVpc Endpoints
Reposted 23 Days AgoSaved
In-Office
Omaha, NE, USA
Mid level
Mid level
Healthtech • Insurance
Owner of enterprise observability and SRE practices: define SLOs/SLA measurement, drive MTTR reduction, lead incident response, maintain service dependency maps and reliability dashboards, and leverage AI/AIOps to automate triage, root cause analysis, and self-healing remediation across vendor and internal platforms.
Top Skills: Ai/AiopsBashChaos EngineeringCi/CdCmdbDashboardingData ModelingDistributed TracingInfrastructure-As-CodeItsm/Ticketing SystemsLog AggregationMonitoring PlatformsObservability PlatformsPowershellPythonSIEMTelemetry
Reposted 23 Days AgoSaved
In-Office
New York, NY, USA
140K-225K Annually
Senior level
140K-225K Annually
Senior level
Fintech
Lead adoption of SRE practices to improve reliability, observability, automation, and incident response. Implement and maintain observability tooling, instrumentation, CI/CD, and infrastructure-as-code. Partner with developers, participate in on-call rotations, drive postmortems, and reduce operational overhead through automation.
Top Skills: AnthropicAWSAws EcsAws EksAzureC#DockerGitlab CiGrafanaLinuxOpenaiPrometheusPuppetPythonSplunkTerraformTypescriptWindows
Reposted 23 Days AgoSaved
In-Office
Bozeman, MT, USA
130K-153K Annually
Senior level
130K-153K Annually
Senior level
Consumer Web • Information Technology • Mobile • Other • Software • App development
Build, maintain, and automate onX's infrastructure platform and deployment pipeline using IaC. Manage Kubernetes/GKE, Terraform, GCP services, observability, and incident response. Improve performance, availability, cost, and developer path to production while participating in on-call rotations and collaborating on architecture.
Top Skills: AirflowBigQueryBigtableChecklyClaude CodeCloud RunCloud SqlCockroachdbGCPGkeGoogle Cloud MonitoringGoogle Cloud StorageGoogle ComposerIamKubernetesNoSQLOpentelemetryOpentofuPrometheusPub/SubRootlySQLTerraform
Reposted 23 Days AgoSaved
In-Office
3 Locations
112K-160K Annually
Senior level
112K-160K Annually
Senior level
Fintech
The Site Reliability Engineer will manage AWS infrastructures, oversee application deployments, and ensure system reliability and security while collaborating with teams.
Top Skills: AWSBashCodebuildCodedeployCodepipelineEc2IamPythonRdsRoute 53S3TerraformVpc
Reposted 23 Days AgoSaved
In-Office or Remote
9 Locations
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
As a Site Reliability Engineer, you'll ensure robotics system reliability, build telemetry integration, and develop tools for diagnostics and automation, collaborating with engineering teams for enhanced production reliability.
Top Skills: C++DatadogGoKubernetesOpentelemetryPrometheusPythonRos2TelegrafTypescript
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted 23 Days AgoSaved
In-Office
Reston, VA, USA
124K-222K Annually
Mid level
124K-222K Annually
Mid level
Cloud • Fintech • HR Tech
Operate and improve service reliability by partnering with development and infrastructure teams. Provide on-call support, troubleshoot distributed systems and cloud infrastructure, implement observability, perform capacity planning and performance optimization, automate operational tasks and infrastructure provisioning, and lead incident management and post-incident analysis.
Reposted 23 Days AgoSaved
Remote
2 Locations
175K-275K Annually
Mid level
175K-275K Annually
Mid level
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills: AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
Reposted 23 Days AgoSaved
Remote or Hybrid
5 Locations
148K-249K Annually
Senior level
148K-249K Annually
Senior level
Transportation
Design and develop Waabi's observability stack, optimize performance, build automation tooling, and support application requirements while leading projects and mentoring teams.
Top Skills: AWSC/C++DockerGoGrafanaJavaKubernetesOpentelemetryPythonRust
Reposted 23 Days AgoSaved
In-Office
Los Angeles, CA, USA
130K-145K Annually
Mid level
130K-145K Annually
Mid level
Events
The Site Reliability Engineer II designs and maintains scalable systems, focusing on automation, monitoring, incident response, and collaboration with developers to enhance operational practices and efficiency.
Top Skills: BashCloud Service OperationsContainersContinuous DeliveryContinuous IntegrationGoInfrastructure As CodeOrchestration PlatformsPython
YesterdaySaved
Remote
4 Locations
Senior level
Senior level
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills: Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
24 Days AgoSaved
In-Office
San Jose, CA, USA
180K-320K Annually
Senior level
180K-320K Annually
Senior level
Software
Lead end-to-end architecture of AI data center networks, DCI, and global backbone for large GPU clusters. Design underlay and overlay networks, IP/VLAN/VXLAN planning, congestion control (PFC/ECN/INT), and optical transport. Produce HLD/LLD, topology diagrams, standards, and SOPs; engage vendors and drive risk mitigation and network improvements.
Top Skills: BgpClosDwdmEcmpEcnEvpnGpu ClustersIn-Band Network Telemetry (Int)InfinibandNvidiaOptical Transport NetworksOspfPfcRocev2Sdn ControllersSegment Routing (Sr-Mpls)Spine-LeafSrv6VlanVxlan
24 Days AgoSaved
In-Office
3 Locations
Senior level
Senior level
Information Technology • Productivity • Software • Manufacturing
Senior individual contributor SRE responsible for reliability, observability, automation, and incident command across multi-account AWS SaaS. Design Terraform modules, CI/CD pipelines, autoscaling/self-healing, OpenTelemetry-based monitoring, and security controls; mentor engineers and integrate validated AI tooling.
Top Skills: Aws Ec2CloudwatchEcs FargateEksGithub ActionsIamLambdaOpenobserveOpentelemetryPagerdutyPythonRds PostgresqlS3TerraformVpc
24 Days AgoSaved
In-Office or Remote
San Mateo, CA, USA
Mid level
Mid level
Cloud • Information Technology
Maintain and improve production service reliability by building automation, monitoring and alerting, participating in on-call incident response, and partnering with engineering and operations to embed reliability practices and reduce operational toil.
Top Skills: AnsibleAWSAzureBashCatchpointCi/CdDockerElkGCPGoGrafanaJenkinsKubernetesPrometheusPythonTerraform
24 Days AgoSaved
In-Office or Remote
San Mateo, CA, USA
Junior
Junior
Cloud • Information Technology
Act as first responder for customer-impacting incidents, monitor and respond to Zabbix alerts, ensure pod and server farm health, run filesystem checks, support Vault deployments and migrations, troubleshoot pod/Ansible/network issues, participate in on-call rotation, document and automate daily tasks, and coordinate escalations with DC techs and management.
Top Skills: AnsibleHashicorp VaultKubernetesLinuxNetworkingZabbix
One Month AgoSaved
In-Office or Remote
San Francisco, CA, USA
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills: Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
Reposted 24 Days AgoSaved
Remote
United States
206K-263K Annually
Expert/Leader
206K-263K Annually
Expert/Leader
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills: Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Reposted 24 Days AgoSaved
In-Office or Remote
2 Locations
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills: AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
2 Days AgoSaved
Hybrid
Bellevue, WA, USA
160K-210K Annually
Senior level
160K-210K Annually
Senior level
Artificial Intelligence • Marketing Tech
Own and scale Cognitiv’s AWS infrastructure while evaluating architecture, networking, security, scalability, and service management. Lead improvements in deployments, monitoring, disaster recovery, and infrastructure-as-code practices. Support co-located Equinix datacenter deployments and hybrid cloud operations alongside a datacenter-focused SRE. Collaborate with engineering and product teams, provide multi-datacenter coverage, and help establish long-term service management best practices.
Top Skills: AnsibleAWSBashDatadogEc2EquinixKubernetesPrometheusPythonTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account