Research Scientist - Robot Learning (VLA / WAM)

Posted 3 Days Ago
Be an Early Applicant
2 Locations
In-Office
Entry level
Artificial Intelligence • Information Technology • Software • Generative AI
The Role
Own end-to-end training of vision-language-action and world-action robot policies, from dataset curation and model architecture through distributed training, evaluation, and deployment. Responsibilities include adapting VLM backbones for control, designing action representations, reducing the sim-to-real gap, building action-conditioned world models, and applying supervised fine-tuning and reinforcement learning. The role also contributes to technical direction for embodied AI research.
Summary Generated by Built In

SpAItial is pioneering the next generation of World Models, pushing the boundaries of generative AI, computer vision, and the simulation of reality. We are moving beyond 2D pixels to build models that natively understand the physics and geometry of our world. Our mission is to redefine how industries, from robotics and AR/VR to gaming and cinema, generate and interact with physically-grounded 3D environments.

We're seeking a Research Scientist to train the policies that turn a world model into a robot that acts. You will own vision-language-action (VLA) and world-action models (WAM) end to end, starting, including data, backbone, action representation, training runs, and the evaluation that tells us whether a policy is genuinely competent or merely lucky. A world model that understands geometry and physics still doesn't act on its own; the policy is what closes that gap. This is a senior, hands-on research role for someone who has already trained manipulation policies that worked, and who can say precisely why the ones that didn't failed.

Responsibilities

  • Own the training pipeline for vision-language-action (VLA) and world-action models (WAM) end to end, from data to a policy running on a robot.

  • Contribute to setting the technical direction for embodied research at SpAItial.

  • Close the sim-to-real gap through domain randomization, system identification, and calibration, and build evaluation that predicts real-world transfer.

  • Adapt VLM backbones for control: encoder choice and adapter strategies, co-training.

  • Curate and weight the training mix across heterogeneous robot datasets, spanning differing embodiments, action spaces, and sensor setups.

  • Design action representation and decoding, including tokenization, chunking, diffusion, and flow-matching action experts.

  • Build the world-model components that predict future observations conditioned on action.

  • Run post-training: supervised fine-tuning onto target embodiments, and RL for robustness beyond demonstrations.

Key Qualifications

  • A PhD in robotics, machine learning, or computer vision with a robot learning focus, from the PhD alone or followed by industry experience.

  • Publications at top venues such as (CoRL, RSS, ICRA, IROS or CVPR, ICCV, ECCV, NeurIPS), open-source work, and/or deployed systems.

  • Deep experience with modern robot policy designs (VLA, WAM, diffusion), trained end to end rather than fine-tuned from a released checkpoint.

  • Strong imitation learning fundamentals, and familiarity with RL fine-tuning of pretrained policies.

  • Fluency with VLM backbones and how to adapt them for control.

  • Expert Python and PyTorch, with multi-node distributed training experience (FSDP or equivalent).

At SpAItial, we are committed to creating a diverse and inclusive workplace. We welcome applications from people of all backgrounds, experiences, and perspectives. We are an equal opportunity employer and ensure all candidates are treated fairly throughout the recruitment process.

Skills Required

  • PhD in robotics, machine learning, or computer vision with a focus on robot learning
  • Publications at leading robotics, computer vision, or machine learning venues, open-source work, and/or deployed systems
  • Deep experience designing and training modern robot policies, including VLA, WAM, or diffusion-based policies end to end
  • Strong imitation learning fundamentals
  • Familiarity with reinforcement learning fine-tuning of pretrained policies
  • Fluency with vision-language model backbones and adapting them for control
  • Expert Python and PyTorch skills
  • Experience with multi-node distributed training using FSDP or equivalent
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
14 Employees
Year Founded: 2024

What We Do

SpAItial is pioneering Spatial Foundation Models (SFMs), a groundbreaking AI paradigm designed to generate and reason about the appearance and physics of real and imagined environments. SFMs possess an intrinsic understanding of space-time, enabling transformative shifts in applications at the intersection of virtual and physical worlds. Unlike existing generative AI technologies such as LLMs, image, or video models, SFMs operate natively in physical space. This significantly advances their cognitive capabilities, which mimics human understanding. SFMs promise to revolutionize various applications across industries, from creating immersive virtual worlds for gaming and entertainment, to advancing CAD engineering and construction, to powering next-generation VR/AR experiences, and enabling sophisticated, physically-intelligent robotics.

Similar Jobs

Celonis Logo Celonis

Associate Applied (AI) Value Engineer (DACH)

Big Data • Information Technology • Productivity • Software • Analytics • Business Intelligence • Consulting
Hybrid
Munich, Bayern, DEU
3000 Employees

Datadog Logo Datadog

Solutions Architect

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
6 Locations
6500 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
München, Bayern, DEU
16000 Employees

Pfizer Logo Pfizer

Digital Operations Agentic Lead - Senior Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
29 Locations
121990 Employees

Similar Companies Hiring

Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account