Principal ML Researcher

Posted Yesterday
Be an Early Applicant
Hiring Remotely in European Union
Remote
Senior level
Artificial Intelligence • Information Technology • Software • Automation
The Role
Lead Toloka’s technical vision and architecture for LLM post-training, reinforcement learning, evaluation, inference optimization, and agentic capabilities. Design scalable GRPO, PPO, DPO, reward modeling, RLAIF, and automated evaluation systems. Mentor senior researchers and engineers, guide cross-functional platform strategy, improve production reliability and cost efficiency, and represent the company through research publications, technical talks, and community engagement.
Summary Generated by Built In

About Toloka

At Toloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety. We have experts in more than 50 different domains—from doctors and lawyers to physicists and engineers—and boast one of the most diverse global crowds, representing over 100 countries and speaking 40+ languages. We are a well-funded startup with an enviable portfolio of clients including Anthropic, Amazon, Microsoft, Poolside, Recraft, and Shopify.

Recently, we secured strategic investment led by Bezos Expeditions and Nebius Group with participation from Mikhail Parakhin, CTO of Shopify and board advisor to leading GenAI companies, who now serves as our Chairman of the Board. Our remote-first team is globally distributed around the world: USA, UK, the Netherlands, Serbia, and more.


About the Team

We are the ML team inside Toloka — we build the machine-learning products that power the platform itself, so every project running on Toloka is faster, cheaper, and more reliable.

A few examples of what we own:

  • Enterprise post-training- we are helping real companies by providing them small LLMs that beat frontier models in quality at a fraction of the cost.
  • Off-the-shelf post-training - we are conducting research on datasets we deliver to clients, showing that training on these datasets will improve their performance.
  • Fine-tuning and RL — adapting frontier and open-source models to Toloka's tasks to hit the right quality at the right cost.
  • Evaluation, benchmarking, cost modeling, and model selection across providers.
  • LLM QA — the core technology behind Toloka's automated quality-check mechanism. Every annotation flowing through Self-Service is reviewed by an LLM agent we design, train, and operate.

We own the full chain. The same team designs the ML solution, ships it to production, keeps it running 24/7, analyzes the results coming back from real projects, and feeds that signal into the next iteration. No hand-off between research, engineering, and operations — it's all us.


About the Position

As a Principal ML Researcher, you will define the overarching technical vision, research strategy, and architecture for Toloka’s core ML and post-training stack. In this high-impact role, you will bridge frontier AI research and large-scale platform engineering.

You will lead technical strategy across greenfield post-training paradigms (such as GRPO, process/outcome reward modeling, and RLAIF), architect resilient automated evaluation ecosystems, and set the standard for how foundation models are adapted and served on our platform. As a technical authority, you will mentor senior ML engineers, collaborate directly with executive leadership, and represent Toloka in the broader AI research community.


What you’ll do

  • Drive Technical Strategy: Define the multi-quarter strategy and architecture for Toloka’s post-training stack, ensuring our fine-tuning, RL, and evaluation capabilities stay ahead of industry trends.
  • Architect Greenfield Post-Training & RL Pipelines: Spearhead the transition beyond basic SFT into advanced RL frameworks (GRPO, PPO/DPO, reward modeling, Process Reward Models, and RLAIF) to power platform-wide LLM alignment and agent behavior.
  • Pioneer Next-Gen Evaluation Harnesses: Design and calibrate bulletproof automated evaluation ecosystems, including human-aligned LLM-as-a-judge frameworks, dynamic benchmarking suites, and regression control setups.
  • Advance Distillation & Inference Optimization: Lead research initiatives on model distillation, quantization, context compression (e.g., gisting), and speculative decoding to optimize latency-cost trade-offs at platform scale.
  • Shape Agentic Platform Capabilities: Architect autonomous guiding agents and tool-use workflows, establishing foundational frameworks for multi-step reasoning, self-correction, and evaluation-driven feedback loops.
  • Technical Leadership & Mentorship: Elevate the engineering and research bar across the team through architectural design reviews, hands-on mentoring of Senior ML Researchers, and establishing best practices for reproducible ML R&D.
  • Cross-Functional & Community Influence: Partner with Product and Engineering directors to translate complex client challenges into scalable platform architecture; author high-impact research write-ups, blog posts, and external tech talks.

What we're looking for

  • 6+ years in ML engineering or applied research, with 3+ years of proven track record leading LLM post-training, fine-tuning, or alignment initiatives at scale.
  • Deep theoretical and practical mastery of LLM alignment: SFT, LoRA/PEFT, RLHF/RLAIF (GRPO, DPO, PPO), reward modeling, and reasoning-oriented post-training.
  • Mastery of LLM Evaluation & Calibration: Proven history of building evaluation pipelines from scratch, addressing LLM-as-a-judge biases, and tightly calibrating automated evals against human gold standards.
  • Expert Distributed Training & Systems Engineering: Strong proficiency in Python, PyTorch, distributed training frameworks (Deepspeed, Megatron, FSDP), and high-performance inference engines (vLLM, TensorRT-LLM).
  • Architectural Vision & Product Leadership: Demonstrated ability to map ambiguous technical challenges into robust platform features, balancing state-of-the-art research with production reliability and cost constraints.
  • Technical Authority: Proven experience acting as a technical leader, mentoring senior researchers/engineers, and influencing cross-functional roadmaps.
  • Language: Fluent spoken and written English (C1), with clear communication skills to articulate complex technical concepts to both internal teams and external clients/community.

What we can offer

  • You will be part of an international, dynamic environment that drives innovation and sets new standards in the AI and technology sector.
  • Competitive compensation package including base salary, bonus, and ESOP.
  • Paid PTO and benefits will vary depending on location.
  • We offer a full remote or hybrid model (if you are based in NL or Serbia).
  • IT setup and home office allowances.

Equal Opportunity Employer:

Toloka is committed to providing equal opportunity and fostering an inclusive environment. We welcome applications from all qualified individuals and do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, marital status, veteran status, disability, or any other characteristic protected by applicable law. Selection decisions are made based on qualifications, merit, and business need.


[Important Notice] Scam Alert Regarding Fake Job Postings

It has come to our attention that an individual or group is fraudulently impersonating Toloka to post fake jobs and solicit personal information from applicants. Please be aware:

  • Official Communication: Our recruiting team will only contact you from an official "toloka.ai" email address. We will NEVER use Gmail, Yahoo, Tolokainc, toloka.inc, or other personal or seemingly business email accounts.
  • Our Process: We will never ask for your bank account details, credit card number, or any fees as part of the application or interview process.
  • Official Listings: All legitimate job openings are posted on our official careers page: https://toloka.ai/careers#job-list

What to do: If you see a suspicious job posting or have been contacted by someone you suspect is a scammer, please do not provide any personal information. Instead, report the incident to us directly at [email protected] and report the profile/post to LinkedIn.We are taking this matter very seriously and are working with the appropriate parties to resolve it.

Thank you for your vigilance!

To learn how we collect, use, disclose, and store personal data, check out our Privacy Notice.


Skills Required

  • 6+ years of experience in machine learning engineering or applied research
  • 3+ years leading LLM post-training, fine-tuning, or alignment initiatives at scale
  • Deep theoretical and practical mastery of LLM alignment, including SFT, LoRA/PEFT, RLHF/RLAIF, GRPO, DPO, PPO, reward modeling, and reasoning-oriented post-training
  • Proven experience building LLM evaluation pipelines and calibrating LLM-as-a-judge systems against human standards
  • Strong proficiency in Python and PyTorch
  • Expertise with distributed training frameworks including DeepSpeed, Megatron, or FSDP
  • Experience with high-performance inference engines such as vLLM or TensorRT-LLM
  • Ability to architect robust ML platform features balancing research innovation, reliability, and cost
  • Experience providing technical leadership, mentoring senior researchers or engineers, and influencing roadmaps
  • Fluent spoken and written English at C1 level
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Amsterdam
1,024 Employees
Year Founded: 2014

What We Do

Toloka empowers businesses to build high quality, safe, and responsible AI. We are the trusted data partner for all stages of AI development from training to evaluation. Toloka has over a decade of experience supporting clients with our unique methodology and optimal combination of machine learning technology and human expertise, offering the highest quality and scalability in the market.

Similar Jobs

Deepgram Logo Deepgram

Solutions Engineer

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Remote
EU
150 Employees

Deepgram Logo Deepgram

Senior Solutions Architect

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Remote
EU
150 Employees

Ruby Labs Logo Ruby Labs

Lead Creative Producer

Information Technology • Software
Remote
7 Locations
28 Employees

Ruby Labs Logo Ruby Labs

Product Manager

Information Technology • Software
In-Office or Remote
15 Locations
28 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account