Researcher - Reinforcement Learning

Reposted One Month Ago
Be an Early Applicant
3 Locations
In-Office
Expert/Leader
Information Technology • Other
The Role
The role involves advancing reinforcement learning techniques for LLMs, including designing training pipelines and evaluating agentic behaviors, contributing to scientific publications.
Summary Generated by Built In

Huawei Canada has an immediate 12-month contract opening for a Reinforcement Learning Researcher.


About the team:

Founded in 2012, the Noah’s Ark lab has evolved into a prominent research organization with notable achievements in academia and industry. The lab’s mission focuses on advancing artificial intelligence and related fields to benefit the company and society. Driven by impactful, long-term projects, the aim is to enhance state-of-the-art research while integrating innovations into the company's products and services, including LLMs, RL, NLP, computer vision, AI theory, and Autonomous driving.

About the job:

  • Enabling Large Language Models (LLMs) to learn from experience, interaction, and environment feedback, moving beyond static fine-tuning toward continual, agentic self-improvement.

  • LLM post-training paradigms (e.g., RLHF, GRPO, reward-free methods, etc.).

  • Agentic reinforcement learning for tool-using and browsing-based LLMs trained in interactive environments.

  • Agentic evaluation and benchmarking, including design of multi-turn, verifiable reasoning tasks.

  • Your work will involve implementing and evaluating new training and evaluation pipelines for reasoning-enhanced LLMs and tool-using agents, scaling experiments on large GPU clusters, and contributing to scientific insights and publications in this emerging area.

About the ideal candidate:

  • PhD degree in Computer Science or related fields or master's degree with comparable experience.

  • Strong foundation in deep learning, including architectures such as Transformers and optimization techniques for large models.

  • Practical or research experience in reinforcement learning, self-supervised learning, or language model fine-tuning.

  • Proven research record in AI by having at least one paper as the first author in top tier venues, such as NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ICRA.

  • Solid proficiency in Python and experience with PyTorch, DeepSpeed, Megatron and other distributed training frameworks.

  • Familiarity with LLM post-training pipelines (RLHF, GRPO/PPO, SFT, LoRA, MoE, etc.) is an asset.

  • Experience with multi-agent RL, tool-use / browser/coding agents, is an asset.

  • Strong communication and writing skills; enthusiasm for open research and collaborative problem-solving.

Informations supplémentaires :

Huawei Canada s'engage à un processus de recrutement équitable, inclusif et accessible. Si vous avez besoin d'un accommodement à n'importe quelle étape du processus d'embauche, veuillez nous en informer, et nous travaillerons avec vous pour répondre à vos besoins.

Toutes les candidatures pour ce poste sont examinées directement par notre équipe de recrutement, nous n'utilisons pas d'outils d'intelligence artificielle pour filtrer ou sélectionner les candidats.

 

Huawei vise à soutenir un environnement de travail francophone pour ses employés au Québec. Nous avons pris des mesures pour éviter de demander une autre langue que le français pour ce poste. Cependant, la maîtrise de l'anglais est essentielle pour ce rôle pour les raisons suivantes :

La personne devra communiquer régulièrement avec des collègues situés en dehors du Québec, où l'anglais est la langue principale utilisée pour la communication entre les bureaux. De plus, la nature des tâches liées à ce poste, qui relève d'un domaine hautement spécialisé de l'intelligence artificielle, nécessite également une connaissance de l'anglais.

L'utilisation du genre masculin a été adoptée afin de faciliter la lecture et n'a aucune intention discriminatoire.

Skills Required

  • PhD in Computer Science or related field
  • Strong foundation in deep learning
  • Practical or research experience in reinforcement learning
  • Proven research record in AI with publications
  • Solid proficiency in Python and distributed training frameworks
  • Familiarity with LLM post-training pipelines
  • Experience with multi-agent RL is an asset
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Sham Chun Hu
1,770 Employees
Year Founded: 1987

What We Do

Founded in 1987, Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices. We are committed to bringing digital to every person, home and organization for a fully connected, intelligent world. We have approximately 197,000 employees and we operate in over 170 countries and regions, serving more than three billion people around the world. In Canada, Huawei conducts innovative and leading edge research in 5G technologies, along with advanced development of emerging cloud, device and network technologies & services. While our renowned Canada Research Centre in the thriving technology landscape of Ottawa, Ontario continues to grow rapidly in size and strategic product initiatives, additional presence has also been established across Canada with R&D facilities in Vancouver, Edmonton, Waterloo, Markham, Montreal, and a R&D office in Quebec City.

Similar Jobs

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Mississauga, ON, CAN
16000 Employees
18-22 Hourly

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Mississauga, ON, CAN
16000 Employees
18-22 Hourly

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Mississauga, ON, CAN
16000 Employees
18-22 Hourly

Block Logo Block

Data Scientist

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
240K-359K Annually

Similar Companies Hiring

Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees
OmniCable Thumbnail
Other
Houston, Texas
815 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account