Site Reliability Engineer

Posted 22 Days Ago
Hiring Remotely in USA
Remote
Senior level
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
The Role
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Summary Generated by Built In
Site Reliability Engineer

Platform and software · shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.

Key responsibilities
  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting

  • Lead incident response and chair post-incident reviews

  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

  • Author and maintain operational runbooks alongside the NOC

  • Manage on-call rotation, escalation paths, and incident-management tooling

  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering

  • Drive chaos engineering, game days, and reliability testing programs

  • Produce SLA performance reports in coordination with the SLA Manager

  • Mentor junior engineers and contribute to engineering culture

Required qualifications
  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications
  • GPU or HPC platform operational experience

  • Familiarity with SLA-driven customer environments and credit calculations

  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)

  • Published SRE content or contributions

Skills Required

  • 5+ years of experience in SRE, DevOps, or production engineering roles
  • Strong programming skills in Go, Python, or both
  • Hands-on experience operating Kubernetes-based platforms at scale
  • Deep familiarity with observability tooling including Prometheus, Grafana, Datadog, and OpenTelemetry
  • Strong incident management experience, including major-incident command
  • GPU or HPC platform operational experience
  • Familiarity with SLA-driven customer environments and credit calculations
  • Experience with chaos engineering tools such as Gremlin or Litmus
  • Published SRE content or contributions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
99 Employees
Year Founded: 2016

What We Do

STN, Inc. is a managed technology and infrastructure provider serving enterprise, regulated, and AI-driven organizations. It designs, operates, and supports secure, scalable systems, including managed IT, cloud and platform services, cybersecurity, data management, compliance engineering, enterprise hardware and software, and GPU One, its GPU-as-a-Service platform for AI training, tuning, inference, and other high-performance workloads. STN emphasizes reliability, audit readiness, and ongoing human support.

Similar Jobs

Optum Logo Optum

Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
92K-164K Annually

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
111K-180K Annually

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
111K-180K Annually

BAE Systems, Inc. Logo BAE Systems, Inc.

Site Reliability Engineer

Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Remote or Hybrid
Merrimack, NH, USA
40000 Employees
118K-201K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account