AI Infrastructure Engineer

Posted 2 Days Ago
2 Locations
In-Office
Junior
Artificial Intelligence • Hardware • Semiconductor • Generative AI
The Role
Build and operate high-performance LLM and multimodal model serving infrastructure using vLLM, SGLang, Kubernetes, and distributed GPU systems. Responsibilities include optimizing parallelism and inference configurations, benchmarking latency and throughput, analyzing model-serving architecture, profiling bottlenecks, and deploying new models to production. The role collaborates with hardware and systems engineers to define backend architecture and service requirements for long-context, agentic, RAG, and coding workloads.
Summary Generated by Built In

About the Role

We're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.

Essential Duties & Responsibilities

  • Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines.

  • Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems.

  • Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization.

  • Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding.

  • Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation.

  • Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime.

  • Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability.

  • Work with model/system engineers to bring newly released models into production efficiently.

  • Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions.

  • Contribute to defining next-generation benchmarks and service requirements as workloads evolve — multi-turn coding, agentic pipelines, RAG, and other long-context use cases.

Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.

  • 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering.

  • Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang.

  • Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution.

  • Strong Python engineering skills.

  • Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs.

  • Clear written and verbal communication skills to work effectively with a small, fully distributed team.

[Preferred Qualifications (optional)]

  • Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, etc.

  • Experience with MoE / long-context model deployment.

  • Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization.

  • Experience with Nsight Systems / PyTorch Profiler.

  • Familiarity with Kubernetes / production GPU serving.

  • Previous startup experience.

Compensation & Benefits

  • Competitive salary with performance-based bonus and early-stage equity grant

  • 100% employer-paid Health, Dental, and Vision coverage for you and your dependents

  • 401(k) match with immediate vesting, and access to financial advisors to help you reach your financial goals

  • 100% employer-paid Life, Disability, and AD&D insurance, plus a fitness stipend and wellness & mental health perks

  • Generous PTO: 20 vacation days, 15 company holidays (including 3 floating days of your choosing)

  • Daily lunch stipend

  • Enterprise-level Claude & ChatGPT access with a generous token budget

  • Well-equipped, sunny offices in Santa Clara, CA & Cambridge, MA with on-site parking and EV charging; on-site fitness center in Santa Clara; gym discounts near our Cambridge office

  • Visa sponsorship and relocation assistance to one of our office hubs

  • A collaborative, continuous-learning environment with smart, dedicated colleagues building the next generation of high-performance computing architecture

The Opportunity

  • Impact: Humanity stands at the dawn of a new industrial revolution driven by AI—one with the potential to redefine how we live on this planet. We are tackling a fundamental challenge at the infrastructure layer: unlocking greater AI capability while dramatically improving efficiency. The work we do here compounds across state-of-the-art AI models, systems, and real-world applications.

  • Timing: Breakthrough technology matters most when it meets the right time. Joining now means real ownership of the company and meaningful influence over product direction and execution. In this early-stage environment, your ideas shape the trajectory of the technology—not just its implementation. You’ll work from first principles, move quickly from insight to execution, and see your contributions directly reflected in what we build.

  • Culture: You’ll work alongside a group of people who care deeply about rigor, clarity, and impact. We value thoughtful disagreement, fast learning, and intellectual fearlessness. This is a place where strong ideas shine, curiosity is encouraged, and growth is a daily practice—not a future promise.

Skills Required

  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience
  • 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering
  • Hands-on experience deploying and performance-tuning vLLM and/or SGLang
  • Strong understanding of LLM inference fundamentals, including prefill versus decode, batching, KV cache, latency and throughput trade-offs, and distributed GPU execution
  • Strong Python engineering skills
  • Working knowledge of continuous batching, KV-cache handling, quantization, and serving SLAs
  • Clear written and verbal communication skills
  • Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, or similar projects
  • Experience with MoE or long-context model deployment
  • Experience with speculative decoding, prefix caching, P/D disaggregation, or attention/KV optimization
  • Experience with Nsight Systems or PyTorch Profiler
  • Familiarity with Kubernetes and production GPU serving
  • Previous startup experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
10 Employees
Year Founded: 2024

What We Do

Netpreme develops memory-to-compute interconnect solutions for generative AI systems, aiming to expand memory bandwidth and capacity, enhance efficiency, and reduce costs.

Similar Jobs

NVIDIA Logo NVIDIA

Infrastructure Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
5 Locations
21960 Employees
184K-357K Annually

EXL Logo EXL

Artificial Intelligence Engineer

Information Technology • Database • Consulting
Remote or Hybrid
United States
30246 Employees
150K-170K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
6 Locations
55000 Employees
209K-286K Annually

CrowdStrike Logo CrowdStrike

Senior Platform Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
140K-215K Annually

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account