Forward Deployed Engineer, Physical AI Infrastructure

Reposted 2 Months Ago
Hiring Remotely in United States
Remote
180K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
The Role
Embedded with strategic customers, own end-to-end design, build, and production rollout of cloud infrastructure for large-scale GPU/HPC AI workloads. Build compute orchestration, platform services, onboarding sandboxes, and reliability/security tooling. Turn field learnings into reusable platform capabilities and partner with Product and Engineering to productize them.
Summary Generated by Built In

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

Nebius is hiring a Forward Deployed Engineer to work alongside Physical AI customers, shape our core infrastructure products, and test ideas for what we should build next.
You will work with enterprises, ISVs, and digital-native companies running training, simulation, synthetic data, evaluation, and model serving workloads. By understanding their systems and helping them reach production, you will uncover requirements that inform our Cloud, Serverless, and Token Factory teams. The goal is to connect needs across Physical AI industries with product decisions, including improvements that benefit enterprise customers more broadly.
You will also investigate opportunities for new Nebius products. You will develop hypotheses about unmet needs, examine demand and existing alternatives, and build prototypes to test with customers. Your findings will help determine which opportunities deserve investment and what an initial product should deliver.
The work is hands-on. You will write code, deploy and benchmark models, build agents, and troubleshoot infrastructure alongside customer engineers.
The position is part of the Physical AI go-to-market team, with close collaboration across customers, NVIDIA, Product, and Engineering.

What you will do
  • Work with customer engineers to understand their workloads, infrastructure constraints, and production requirements.
  • Translate Physical AI use cases into requirements that Cloud, Serverless, and Token Factory teams can use, backed by measurements and customer evidence.
  • Identify enterprise adoption barriers and common needs across industries. Work with Product and Engineering to assess gaps and test improvements with customers.
  • Build agents that set up, run, and debug Physical AI workloads, connecting them to clusters, schedulers, storage, and model servers.
  • Deploy and benchmark open-weight models with serving stacks such as vLLM or SGLang, then tune them for customer workloads.
  • Build prototypes and run technical evaluations to answer specific questions about performance, reliability, and product fit.
  • Assess demand and existing alternatives for opportunities outside current product teams’ scope, then build where there is a clear market.
  • Debug across Python, agents, model servers, Kubernetes, GPUs, networking, and storage. Turn successful setups into reusable scripts, deployments, and short guides.
What we are looking for
  • Three or more years in software engineering, cloud infrastructure, DevOps, ML infrastructure, or similar, including ownership of a production service, cluster, or platform.
  • Experience working directly with customers or external engineers to scope problems, run evaluations, and explain technical tradeoffs.
  • Ability to turn an unfamiliar customer problem into concrete infrastructure requirements.
  • Practical experience building with LLMs or agents, at work or independently, with an understanding of their internals (e.g., transformers, KV caching, or tool calling) and how to evaluate their outputs.
  • Familiarity with enterprise infrastructure requirements, including identity and access management, network isolation, security, monitoring, and reliability.
  • Strong Python and Linux skills: you can read unfamiliar code, automate a setup, and debug from logs and metrics.
  • Practical experience with containers and Kubernetes, including configuration, networking, storage, and scheduling.
  • A record of getting useful software working with incomplete requirements, measuring results, and following through with the people using it.
Helpful experience
  • Enterprise customer engineering, solutions architecture, consulting, or taking a new product from an initial customer problem to adoption.
  • Distributed compute such as Ray, Slurm, Spark, or multi-GPU jobs.
  • NVIDIA GPU workloads, including memory management, drivers, CUDA, NCCL, and monitoring.
  • Model serving with vLLM, SGLang, TensorRT-LLM, or TGI.
  • Robotics, simulation, synthetic data, reinforcement learning, world models, or tools such as Isaac Sim, Isaac Lab, Omniverse, and Cosmos.
  • Infrastructure automation and enterprise integrations, including Terraform, CI/CD, private connectivity, and identity federation.

Pay Transparency

We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.

Base Compensation Range
$179,500—$224,300 USD

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI 

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. 

If you need accommodations during the application process, please let us know.

Skills Required

  • 6+ years of hands-on backend, cloud infrastructure, platform engineering, or SRE experience
  • At least 2 years in a customer-facing or deployment-oriented technical role (FDE, founding engineer, tech lead embedded with customers, or equivalent)
  • Experience building distributed systems, job orchestration, compute platforms, internal developer platforms, or ML infrastructure
  • Strong systems/backend programming skills (Python, Go, or similar)
  • Fluency using modern AI coding tools (e.g., Claude Code, Codex, Cursor) for rapid development
  • Experience with Kubernetes, containers (Docker), CI/CD, observability, cloud networking, cloud storage, IAM/RBAC, and infrastructure as code
  • Familiarity with GPU and HPC workloads including batch, training, and inference pipelines
  • Proven ability to debug infrastructure issues across application, network, storage, compute, and orchestration layers
  • Strong security and reliability instincts (isolation, RBAC, uptime, traceability)
  • High agency with strong written and verbal communication
  • Authorized to work in the country in which you apply and able to provide proof of employment eligibility
  • Prior experience as a Forward Deployed Engineer or equivalent customer-embedded engineering function
  • Experience with Nebius, AWS, GCP, Azure, or Lambda Labs
  • Experience with Slurm, Soperator, Kubernetes GPU scheduling, Ray, Argo, Airflow, Metaflow or similar orchestration tools
  • Experience with ML training infrastructure, model serving, simulation workloads, or large-scale data pipelines
  • Experience supporting enterprise customers, design partners, or production pilots
  • Familiarity with NVIDIA GPU infrastructure, CUDA workloads, Isaac Sim, Omniverse, or simulation-at-scale

Nebius Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Nebius and has not been reviewed or approved by Nebius.

  • Fair & Transparent Compensation — Pay is considered competitive or good across many roles, with overall sentiment leaning positive. While experiences vary by team and location, compensation frequently stands out as a strength.
  • Flexible Benefits — Work arrangements are described as remote-friendly with hybrid options and flexibility across time zones. Some roles reference home-office stipends that support distributed work.
  • Healthcare Strength — Company materials and job postings describe comprehensive employer-paid medical, dental, and vision coverage for employees and families. Details may vary by country and entity.

Nebius Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Amsterdam
473 Employees

What We Do

Cloud platform specifically designed to train AI models

Similar Jobs

Arcadia Logo Arcadia

Principal Analyst, Medicaid Policy Analysis

Big Data • Fitness • Healthtech • Information Technology • Software • Analytics
Remote
USA
360 Employees
190K-215K Annually

Corporate Tools LLC Logo Corporate Tools LLC

Sr. iOS App Developer

eCommerce • Legal Tech • Professional Services • Software • Data Privacy
Remote or Hybrid
Post Falls, Idaho, USA
1200 Employees
150K-150K Annually

Wells Fargo Logo Wells Fargo

Branch Manager South Jax St Johns

Fintech • Financial Services
Remote or Hybrid
Ponte Vedra, FL, USA
205000 Employees

Headway Logo Headway

Senior Product Manager

Consumer Web • Healthtech • Professional Services • Social Impact • Software
Remote
USA
819 Employees
223K-279K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
65 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account