Software Engineer - Training/Inference (C++)

Reposted 2 Months Ago
Be an Early Applicant
2 Locations
In-Office
180K-440K Annually
Senior level
Information Technology
The Role
Design and optimize large-scale, low-latency model serving systems end-to-end. Implement distributed infrastructure (batching, caching, load balancing, autoscaling), accelerate GPU inference (kernels, quantization, codegen), build reliable high-concurrency serving, and create tooling and CI/CD for deployment and benchmarking.
Summary Generated by Built In

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

About the Role:
  • We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability.
  • As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end. You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency).
  • This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale

Responsibilities: 

  • Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
  • Optimize latency and throughput of model inference under real production workloads.
  • Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency.
  • Benchmark, fine-tune, and accelerate inference engines (including low-level GPU kernel work and code generation).
  • Develop custom tools to trace, replay, and fix issues across the full stack — from orchestration down to GPU kernels.
  • Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.
  • Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.
BASIC QUALIFICATIONS:
  • Deep low-level systems programming (C/C++ or Rust)
  • Experience with large-scale, high-concurrent production serving.
  • Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
  • Strong background in system optimizations: batching, caching, load balancing, parallelism.
  • Low-level inference optimizations: GPU kernels, code generation.
  • Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics.
  • Experience with testing, benchmarking, and reliability of inference services.
  • Experience designing and implementing CI/CD infrastructure for inference.
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Skills Required

  • Deep low-level systems programming (C/C++ or Rust)
  • Experience with large-scale, high-concurrent production serving
  • Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM)
  • Strong background in system optimizations: batching, caching, load balancing, parallelism
  • Low-level inference optimizations: GPU kernels, code generation
  • Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics
  • Experience with testing, benchmarking, and reliability of inference services
  • Experience designing and implementing CI/CD infrastructure for inference
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
96 Employees

What We Do

Understand the Universe

Similar Jobs

Wipfli Logo Wipfli

Manager, Accounting Advisory - Behavioral Health Industry

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
2900 Employees
107K-160K Annually

Wipfli Logo Wipfli

Senior Analyst FP&A

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
2900 Employees
80K-108K Annually

Boeing Logo Boeing

EA-18G Growler Electrician

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Oak Harbor, WA, USA
170000 Employees
73K-97K Annually

Boeing Logo Boeing

Intellectual Property Release Specialist

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Hybrid
Seattle, WA, USA
170000 Employees
99K-155K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account