Job Description:
DataRobot delivers AI that maximizes impact and minimizes business risk. Our platform and applications integrate into core business processes so teams can develop, deliver, and govern AI at scale. DataRobot empowers practitioners to deliver predictive and generative AI, and enables leaders to secure their AI assets. Organizations worldwide rely on DataRobot for AI that makes sense for their business — today and in the future.
This effort is about building a multi-target joint probabilistic foundation model that can be used across industries to tackle some of the hardest real-world problems. The ambition goes beyond what tabular and time-series foundation models usually do: one model should support temporal forecasting, unordered tabular regression and classification, missing-data completion, and mixed-modality inputs and outputs while learning coherent joint structure across connected variables, rows, horizons, and scenarios. The primary objective is to create real value in serious use cases rather than optimize benchmark scores in isolation, although strong results on standard benchmarks are expected to follow closely.
This role is designed as a research-engineering internship for someone operating at post-doctoral level, or very close to it, who wants to shape the architecture of a probabilistic foundation model. The center of gravity is the model itself: the encoders that read temporal, tabular, and mixed-modality inputs, the attention and sequence mechanisms that carry structure across variables, rows, and horizons, and the distributional output heads and decoding schemes that turn hidden states into coherent joint predictions. We are looking for someone who can reason precisely about what an architecture can and cannot represent, and who then implements, trains, and evaluates the resulting PyTorch models in a production-quality codebase. A background in stochastic modeling is a strong plus rather than a prerequisite: the model is trained on synthetic data drawn from stochastic dynamics, so a candidate who can also reason about stochastic differential equations and design synthetic data generators can contribute on both sides of the training loop.
Required Skills:
- Strong PyTorch skills and hands-on experience building and training deep models.
- Solid understanding of Transformers, attention variants, and long-context or state-space sequence architectures, including the tradeoffs between quality, memory, and latency.
- Experience with probabilistic modeling in neural networks: distributional output heads, likelihood-based or proper-scoring-rule losses, mixture or flow models, or related uncertainty-aware learning setups.
- Strong foundation in probability and statistics, enough to reason about densities and masses, dependence between variables, and calibration.
- Ability to connect architectural ideas to working GPU-native implementations, controlled experiments, and diagnostics that show why a change helped.
- Strong engineering habits: readable code, tests, reproducible experiments, and disciplined evaluation of model changes.
- Ability to debug training instability, reason about why results changed, and iterate quickly from hypothesis to evidence.
Strong Plus:
- Working knowledge of stochastic processes and stochastic differential equations: drift and diffusion, jumps, regime switching, mean reversion, heavy tails, and how such dynamics are simulated numerically.
- Experience designing synthetic data generators or simulation-based training curricula, and an understanding of how the training distribution shapes what a model learns.
- Depth in a domain with rich stochastic structure such as finance, energy, commodities, or a similarly quantitative field.
- Familiarity with tabular or mixed-modality deep learning: categorical, ordinal, count, and bounded targets alongside continuous ones.
What You Can Expect To Work On:
- Designing and improving the core model architecture: input encoders for temporal, unordered tabular, and mixed-modality data; attention factorizations across variables, rows, and horizons; the distributional output heads and decoding strategies that produce coherent joint samples.
- Running controlled architecture studies, from ablations and scaling behavior to memory and throughput profiles, and turning the results into design decisions.
- Building scalable PyTorch implementations that support larger input and output spaces, better throughput, and tighter memory budgets.
- Studying how architectural choices, data-generation choices, and inference constraints interact to change benchmark quality and real-world usefulness.
- For candidates with the stochastic-modeling background: extending the synthetic-data engine with richer stochastic dynamics, constraints, dependence structures, heavy tails, and regime behavior.
- Turning research ideas into robust implementations and credible empirical results.
What You Should Expect:
- Work that sits close to the core of the project rather than at the edges.
- A fast research loop where good ideas can move quickly from hypothesis to implementation to benchmark.
- A real chance to contribute to research that aims for top-tier, state-of-the-art outcomes rather than only incremental internal work.
- A team that values rigorous thinking, clean code, and practical usefulness at the same time.
- Exposure to problems that combine deep learning, stochastic modeling, and real deployment constraints instead of isolating only one of those dimensions.
- A role that fits someone with broad interests who wants to work across both advanced modeling and serious implementation work rather than staying narrowly specialized.
The talent and dedication of our employees are at the core of DataRobot’s journey to be an iconic company. We strive to attract and retain the best talent by providing competitive pay and benefits with our employees’ well-being at the core. Here’s what your benefits package may include depending on your location and local legal requirements: Medical, Dental & Vision Insurance, Flexible Time Off Program, Paid Holidays, Paid Parental Leave, Global Employee Assistance Program (EAP) and more!
DataRobot Operating Principles:
- Wow Our Customers
- Set High Standards
- Be Better Than Yesterday
- Be Rigorous
- Assume Positive Intent
- Have the Tough Conversations
- Be Better Together
- Debate, Decide, Commit
- Deliver Results
- Overcommunicate
Research shows that many women only apply to jobs when they meet 100% of the qualifications while many men apply to jobs when they meet 60%. At DataRobot we encourage ALL candidates, especially women, people of color, LGBTQ+ identifying people, differently abled, and other people from marginalized groups to apply to our jobs, even if you do not check every box. We’d love to have a conversation with you and see if you might be a great fit.
DataRobot is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. DataRobot is committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. Please see the United States Department of Labor’s EEO poster and EEO poster supplement for additional information.
Use of Artificial Intelligence in Our Hiring Process
DataRobot uses approved AI-powered tools to support the hiring process in selected regions. These tools may assist in writing job descriptions, reviewing applications, assessing qualifications, and evaluating candidate materials. All decisions regarding applications are made by members of the DataRobot team.
All applicant data submitted is handled in accordance with our Applicant Privacy Policy.
Skills Required
- Strong PyTorch skills and hands-on experience building and training deep models
- Solid understanding of Transformers, attention variants, and long-context or state-space sequence architectures
- Experience with probabilistic modeling in neural networks, including distributional output heads, likelihood-based or proper-scoring-rule losses, mixture or flow models, or related uncertainty-aware learning
- Strong foundation in probability and statistics, including densities, masses, variable dependence, and calibration
- Ability to connect architectural ideas to GPU-native implementations, controlled experiments, and diagnostic analysis
- Strong engineering habits, including readable code, tests, reproducible experiments, and disciplined evaluation
- Ability to debug training instability, explain result changes, and iterate from hypothesis to evidence
- Working knowledge of stochastic processes and stochastic differential equations
- Experience designing synthetic data generators or simulation-based training curricula
- Depth in finance, energy, commodities, or another quantitative domain with rich stochastic structure
- Familiarity with tabular or mixed-modality deep learning
DataRobot Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about DataRobot and has not been reviewed or approved by DataRobot.
-
Fair & Transparent Compensation — Pay is described as competitive and often positioned as fair relative to comparable roles, with salary bands that can reach the upper end for certain positions. Overall compensation is also framed as strong enough to attract and retain top talent.
-
Equity Value & Accessibility — Equity is consistently included as part of the compensation package, with restricted stock awards and stock-related programs featured as meaningful components. The presence of equity for employees is treated as a key differentiator in total rewards.
-
Healthcare Strength — Health, dental, and vision coverage are presented as comprehensive, and the overall benefits bundle is portrayed as strong. Additional coverage like pet insurance further reinforces the breadth of healthcare-related support.
DataRobot Insights
What We Do
DataRobot is the AI Cloud leader, delivering a unified platform for all users, all data types, and all environments to accelerate delivery of AI to production. Trusted by global customers across industries and verticals, including a third of the Fortune 50, delivering over a trillion predictions for leading companies globally.









