Forward Deployed Engineer

Posted Yesterday
Be an Early Applicant
Palo Alto, CA, USA
In-Office
150K-195K Annually
Senior level
Artificial Intelligence • Machine Learning • Software
The Role
Customer-facing engineer who owns technical wins from POC to production: designs and runs reproducible inference benchmarks, tunes model-to-hardware deployments, conducts vendor bake-offs, handles security/compliance, drives post-signature account health, and produces reusable benchmark reports and reference architectures to accelerate sales and deployments.
Summary Generated by Built In
About DeepInfra

DeepInfra is building the foundation for companies to use modern AI in production — simply, reliably, and at scale. Our team has deep experience building large systems that serve hundreds of millions of users, and we're bringing that same level of rigor to a rapidly evolving AI inference space. Our mission is to make advanced AI available to people and teams everywhere.

We're an early, tight-knit team where you can influence product direction, try bold ideas, and drive meaningful work forward quickly. If you want to join a fast-growing company at a defining moment, we'd love to talk.

DeepInfra is backed by leading investors including A.Capital, Felicis, 500 Global, Georges Harik, Samsung Next, Supermicro, Upper90, Peak6, SVAngel and Nvidia.

Why this role matters

As DeepInfra's enterprise pipeline grows, our customers need a technical partner who can run rigorous evals, defend benchmarks, and speak fluently to both engineering and procurement — someone who can own the technical win from first call through production.

This is a pioneering role. You'll work closely with Sales, our co-founders, and the engineering team on the deals that matter most. You'll own the technical win end to end: running head-to-head bake-offs against leading AI providers, tuning deployments on the latest hardware, and turning what you learn into reusable assets that make every future deal faster to close. As our first FDE, you'll also define what the function looks like as GTM scales.

What You'll Do

  • Own the technical win and the POC timeline, working closely with Sales and Engineering, from call one.
  • Design and run reproducible benchmark harnesses (TTFT, ITL, throughput/GPU, p95/p99) and quality-parity evals.
  • Run head-to-head bake-offs against leading AI providers — and win them.
  • Tune model-to-hardware deployments on B200/B300/GB300 NVL72.
  • Build cost-per-token models and write migration plans.
  • Handle enterprise security and compliance review, and get deployments to launch readiness.
  • Own account health post-signature, driving usage reviews and expansion.
  • Turn what you learn into reusable benchmark reports, reference architectures, and AE enablement material.

What You Bring

  • Customer-facing engineering with an owned technical outcome at an infrastructure or ML platform company.
  • Strong Python skills.
  • Dual-audience presence with commercial instinct — credible with a skeptical staff engineer, clear with a CFO, and able to tell a technical objection from a procurement one.

Bonus

  • Hands-on experience with inference internals: vLLM, SGLang, or TRT-LLM, batching, KV cache math, quantization.
  • Experience with agentic or coding-assistant workloads at scale.
  • Prefix-cache-heavy long context workloads.
  • Diffusion image/video, ASR/TTS, or multi-LoRA serving.
  • Open-source contributions to vLLM or SGLang.
  • Deep NVLink/InfiniBand topology knowledge.

Why DeepInfra

  • Define DeepInfra's Forward Deployed Engineering function from day one and have a direct impact on its direction.
  • Work directly with co-founders and the inference team on the deals that matter most.
  • Join a small, high-performing team where your work ships quickly and reaches customers around the world.
  • Help shape how enterprises adopt some of the world's leading open-source AI models.

How we work

Three traits define the people who thrive here, and this role leans on all three.

Initiative. We take ownership and step in where we can add value. Whether it’s starting something new, improving what exists, or helping move ideas forward, we aim to be proactive and thoughtful in how we contribute.

Drive. We’re energized by hard problems. Building AI infrastructure is complex, and we lean into that. We care about doing things well, moving fast, and continuously improving — because solving meaningful challenges is what motivates us.

Grit. Things don’t always work on the first try — and that’s expected. We stay persistent, adapt quickly, and learn as we go. We take setbacks seriously, but not personally, and use them to get better.
Compensation
The base pay range for this role is $150,000 – $195,000 per year.

Skills Required

  • Customer-facing engineering with an owned technical outcome at an infrastructure or ML platform company.
  • Strong Python skills.
  • Dual-audience presence with commercial instinct (credible with engineers and CFO/procurement).
  • Experience tuning model-to-hardware deployments (B200/B300/GB300/NVL72) and GPU inference optimization.
  • Hands-on knowledge of inference internals (vLLM, SGLang, TRT-LLM), batching, KV cache math, and quantization.
  • Experience with agentic or coding-assistant workloads, long-context/prefix-cache heavy workloads, or multi-LoRA serving.
  • Experience with diffusion image/video, ASR/TTS workloads at scale.
  • Open-source contributions to vLLM or SGLang or deep NVLink/InfiniBand topology knowledge.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, California
20 Employees
Year Founded: 2022

What We Do

Let Deep Infra run your ML infrastructure. Just use our top AI models using a simple API or deploy your own model with us.

Similar Jobs

CoreWeave Logo CoreWeave

Forward Deployed Engineer, AI Agents

Cloud • Information Technology • Machine Learning
In-Office
2 Locations
1450 Employees
182K-242K Annually

Granica Logo Granica

Forward Deployed Engineer

Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
In-Office
Mountain View, CA, USA
45 Employees
160K-220K Annually

SailPoint Logo SailPoint

Forward Deployed Engineer

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
United States
2461 Employees
106K-179K Annually

Shield AI Logo Shield AI

Forward Deployed Engineer, Operations & Sustainment (R5497)

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
In-Office or Remote
3 Locations
150K-230K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account