Platform and software · shared across customers
Reports to: Director, Site Reliability
Location: Remote (US)
Department: Cloud Platform Engineering / SRE/Reliability
Position summaryThe Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.
Key responsibilitiesDefine and operate Service Level Objectives (SLOs) aligned with customer SLAs
Build and maintain the observability stack including metrics, logs, traces, and alerting
Lead incident response and chair post-incident reviews
Drive automation to reduce toil and improve mean-time-to-recover (MTTR)
Author and maintain operational runbooks alongside the NOC
Manage on-call rotation, escalation paths, and incident-management tooling
Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering
Drive chaos engineering, game days, and reliability testing programs
Produce SLA performance reports in coordination with the SLA Manager
Mentor junior engineers and contribute to engineering culture
5+ years in SRE, DevOps, or production engineering roles
Strong programming skills in Go, Python, or both
Hands-on experience operating Kubernetes-based platforms at scale
Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
Strong incident management experience including major-incident command
GPU or HPC platform operational experience
Familiarity with SLA-driven customer environments and credit calculations
Experience with chaos engineering tools (Gremlin, Litmus, or similar)
Published SRE content or contributions
Skills Required
- 5+ years of experience in SRE, DevOps, or production engineering roles
- Strong programming skills in Go, Python, or both
- Hands-on experience operating Kubernetes-based platforms at scale
- Deep familiarity with observability tooling including Prometheus, Grafana, Datadog, and OpenTelemetry
- Strong incident management experience, including major-incident command
- GPU or HPC platform operational experience
- Familiarity with SLA-driven customer environments and credit calculations
- Experience with chaos engineering tools such as Gremlin or Litmus
- Published SRE content or contributions
What We Do
STN, Inc. is a managed technology and infrastructure provider serving enterprise, regulated, and AI-driven organizations. It designs, operates, and supports secure, scalable systems, including managed IT, cloud and platform services, cybersecurity, data management, compliance engineering, enterprise hardware and software, and GPU One, its GPU-as-a-Service platform for AI training, tuning, inference, and other high-performance workloads. STN emphasizes reliability, audit readiness, and ongoing human support.

.png)





