Infrastructure operations · shared across customers
Reports to: Manager, NOC (or Director, Service Operations)
Location: Remote (US) with assigned shift; rotating coverage
Department: Infrastructure & DC Operations / Network Engineering
Position summaryThe NOC Engineer operates STN's 24/7 monitoring and first-response capability for GPU One (GPUaaS) infrastructure. The role triages alerts, executes documented runbooks, and coordinates with on-call specialists during incidents to protect customer SLAs.
Key responsibilitiesMonitor infrastructure alerts, customer SLA dashboards, and system health on a 24/7 basis
Triage incidents and engage on-call SREs, Network, Hardware, or Field Engineering as needed
Execute documented runbooks for common platform, network, and hardware issues
Manage the incident lifecycle including initial customer notification and status updates
Coordinate planned maintenance windows and change windows with internal teams and customers
Update status pages and customer-facing communications during incidents
Maintain shift handoff documentation and active-incident logs
Support ticket queue handling including Tier 1 ticket resolution
Contribute to continuous improvement of monitoring coverage, alert quality, and runbooks
Work rotating shifts including nights, weekends, and holidays
3+ years in a NOC, SOC, or IT operations function
Hands-on experience with monitoring tools (Datadog, Prometheus, Grafana, PagerDuty, or equivalent)
Strong Linux and basic networking fundamentals
Excellent written and verbal communication, particularly under pressure
Willingness and ability to work rotating shifts including overnight coverage
GPU, HPC, or large-scale cloud infrastructure background
ITIL Foundations certification
Demonstrated on-call and major-incident response experience
Scripting skills (Python, Bash) for runbook automation
Skills Required
- 3+ years of experience in a NOC, SOC, or IT operations function
- Hands-on experience with monitoring tools such as Datadog, Prometheus, Grafana, PagerDuty, or equivalent
- Strong Linux fundamentals
- Basic networking fundamentals
- Excellent written and verbal communication, particularly under pressure
- Willingness and ability to work rotating shifts, including overnight coverage
- GPU, HPC, or large-scale cloud infrastructure experience
- ITIL Foundations certification
- On-call and major-incident response experience
- Python or Bash scripting skills for runbook automation
What We Do
STN, Inc. is a managed technology and infrastructure provider serving enterprise, regulated, and AI-driven organizations. It designs, operates, and supports secure, scalable systems, including managed IT, cloud and platform services, cybersecurity, data management, compliance engineering, enterprise hardware and software, and GPU One, its GPU-as-a-Service platform for AI training, tuning, inference, and other high-performance workloads. STN emphasizes reliability, audit readiness, and ongoing human support.








