Senior Platform Engineer

Posted Yesterday
Hiring Remotely in USA
Remote
Senior level
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
The Role
Build and operate the multi-tenant platform layer for a GPU cloud service. Responsibilities include Kubernetes-based orchestration, tenant isolation, platform APIs and tools, GPU cluster automation, infrastructure as code, reliability and observability partnerships, customer issue support, environment templates, rollout tooling, architecture reviews, and reducing operational toil.
Summary Generated by Built In
Senior Platform Engineer

Platform and software · shared across customers

Reports to: Director, Platform Engineering (or Chief Architect)

Location: Remote (US) or Pleasanton, CA (hybrid)

Department: Cloud Platform Engineering / GPU Platform Engineering

Position summary

The Senior Platform Engineer builds and operates the multi-tenant orchestration, scheduling, and customer-facing platform layer that turns raw GPU infrastructure into a usable cloud service. This role is the software backbone of GPU One (GPUaaS).

Key responsibilities
  • Design and build the orchestration layer (Kubernetes, Slurm, Run:ai, or comparable)

  • Manage multi-tenant isolation including namespaces, networking, storage, and quotas

  • Build customer-facing platform APIs, CLIs, web portals, and SDKs

  • Implement and operate image management, GPU operator, and node provisioning automation

  • Drive infrastructure-as-code and automation across the platform stack

  • Partner with SRE on platform reliability, SLO definition, and observability

  • Support TAM and Support engineers on customer-impacting platform issues

  • Maintain customer environment templates, configuration management, and rollout tooling

  • Participate in architecture review, design discussions, and technical roadmap

  • Drive continuous platform improvement and reduce operational toil

Required qualifications
  • 6+ years in platform engineering, SRE, or cloud engineering at scale

  • Deep Kubernetes expertise including CRDs, operators, and multi-tenant patterns

  • Strong programming skills in Go, Python, or both

  • Experience operating GPU clusters or AI infrastructure at production scale

  • Bachelor's degree in computer science or equivalent experience

Preferred qualifications
  • Experience with NVIDIA GPU Operator, MIG, MPS, and NCCL operator patterns

  • Familiarity with Slurm operator, Run:ai, KubeRay, or comparable AI orchestration

  • Service mesh experience (Istio, Linkerd) and multi-cluster networking

  • Open source contributions in the cloud-native or AI infrastructure ecosystem

Skills Required

  • 6+ years of experience in platform engineering, SRE, or cloud engineering at scale
  • Deep Kubernetes expertise, including CRDs, operators, and multi-tenant patterns
  • Strong programming skills in Go, Python, or both
  • Experience operating GPU clusters or AI infrastructure at production scale
  • Bachelor's degree in computer science or equivalent experience
  • Experience with NVIDIA GPU Operator, MIG, MPS, and NCCL operator patterns
  • Familiarity with Slurm operator, Run:ai, KubeRay, or comparable AI orchestration
  • Service mesh experience with Istio, Linkerd, or similar and multi-cluster networking
  • Open source contributions in the cloud-native or AI infrastructure ecosystem
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
99 Employees
Year Founded: 2016

What We Do

STN, Inc. is a managed technology and infrastructure provider serving enterprise, regulated, and AI-driven organizations. It designs, operates, and supports secure, scalable systems, including managed IT, cloud and platform services, cybersecurity, data management, compliance engineering, enterprise hardware and software, and GPU One, its GPU-as-a-Service platform for AI training, tuning, inference, and other high-performance workloads. STN emphasizes reliability, audit readiness, and ongoing human support.

Similar Jobs

CrowdStrike Logo CrowdStrike

Senior Platform Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

inKind Logo inKind

Senior Platform Engineer

eCommerce • Fintech • Food • Mobile • Social Impact
Remote or Hybrid
USA
170 Employees
190K-200K Annually

Webflow Logo Webflow

Senior Platform Engineer

Artificial Intelligence • Enterprise Web • Software • Design • Generative AI
Easy Apply
Remote
U.S.
800 Employees
187K-255K Annually

GitLab Logo GitLab

Senior Software Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
4 Locations
2500 Employees
139K-235K Annually

Similar Companies Hiring

Rain Thumbnail
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
New York, NY
100 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account