ML Researcher - Image / Video Diffusion

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
In-Office
Entry level
Artificial Intelligence • Software • Design • Generative AI
The Role
Research and engineer large-scale image and video diffusion models. Train and optimize distributed models on GPU clusters, implement parallelism strategies, profile and debug training infrastructure, improve data and model quality, design experiments and evaluations, and develop fault-tolerant solutions for hardware and numerical issues. The role requires rapid iteration, strong research judgment, custom data pipeline development, and hands-on expertise with PyTorch, distributed training, low-precision computation, and diffusion training across pretraining, preference optimization, and reinforcement learning.
Summary Generated by Built In
About Krea

At Krea, we are building next-generation AI creative tools.

We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2, our first foundation model, built completely from scratch for aesthetic diversity and stylistic control.

We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.

 

We're looking for an experienced Researcher with engineering skills who can work on large-scale image and video models training experiments, with experience training image models at scale.

Our culture
  • We work full-time and in-person at our North Beach office in San Francisco.

  • We believe that demonstrated interest in the creative space is key: our team includes musicians, designers, visual artists and more.

  • Fast iteration and execution speed. Bias towards action, agency, and independence.

What you'll do
  • Train diffusion models for image and video generation on large GPU clusters.

  • Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.

  • Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.

  • Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.

  • Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.

  • Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.

What we're looking for
  • Proven track record in working with image or video models at scale (publications or open-source contributions a plus).

  • Strong proficiency in PyTorch and understanding of its inner workings.

  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.

  • Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements.

  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.

  • Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.

  • Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.

  • Being comfortable working in a goal-oriented research environment.

  • Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.

  • Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items.

  • Good research taste — bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.

  • Ability to iterate rapidly, and propose creative research directions.

  • Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.

What we offer
  • Team: Work alongside a world-class team building the future of AI creative tooling

  • Impact: Significant scope and company-wide impact

  • Competitive compensation: generous salary & equity packages

  • Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage

  • Time off: Flexible PTO policy

  • Financial planning: 401k with a 4% company-sponsored match

  • Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it

  • Transit: Ubers covered to & from the office

  • Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).

  • And more!

Please note the above benefits & perks are for full-time employees

Skills Required

  • Proven experience working with image or video models at scale
  • Strong proficiency in PyTorch and understanding of its inner workings
  • Strong background in distributed training paradigms including FSDP, CP, SP, USP, TP, and EP
  • Experience profiling and debugging large distributed training runs
  • Knowledge of low-precision training and inference using FP8, NVFP4, and MXFP8
  • Understanding of diffusion model training across pretraining, midtraining, preference optimization, and reinforcement learning
  • Ability to work with underspecified goals and translate them into concrete plans and experiments
  • Ability to design custom data pipelines and improve data quality
  • Publications or open-source contributions in image or video modeling
  • Knowledge of related research fields including LLMs, VLMs, representation learning, and robotics
  • Demonstrated interest in creative fields
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
1,019 Employees
Year Founded: 2007

What We Do

Krea is a generative AI creative platform offering AI tools for creatives to generate, edit, and enhance images, video, and 3D content. It uses artificial intelligence to generate visuals tailored to unique styles, concepts, or products.

Similar Jobs

Snap Inc. Logo Snap Inc.

Hardware Engineer

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Hybrid
2 Locations
5000 Employees
195K-343K Annually

SOPHiA GENETICS Logo SOPHiA GENETICS

Scientist

Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Remote or Hybrid
3 Locations
450 Employees

PNC Bank Logo PNC Bank

Software Engineer

Machine Learning • Payments • Security • Software • Financial Services
Remote or Hybrid
USA
55000 Employees
75K-150K Annually

PNC Bank Logo PNC Bank

Senior Software Engineer

Machine Learning • Payments • Security • Software • Financial Services
Remote or Hybrid
USA
55000 Employees

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account