ML/AI Engineer

Reposted 7 Days Ago
Be an Early Applicant
Amsterdam
In-Office
Mid level
Artificial Intelligence • Information Technology • Consulting
The Role
The ML/AI Engineer will benchmark GPU platforms for AI workloads, analyze GPU performance, optimize ML workloads, and develop performance visualization tools.
Summary Generated by Built In

Why work at Nebius
Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field.

Where we work
Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 800 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team.

About Nebius AI 

Nebius AI is an AI cloud platform with one of the largest GPU capacities in Europe. Launched in November 2023, the Nebius AI platform provides high-end, training-optimized infrastructure for AI practitioners. As an NVIDIA preferred cloud service provider, Nebius AI offers a variety of NVIDIA GPUs for training and inference, as well as a set of tools for efficient multi-node training. 

Nebius AI owns a data center in Finland, built from the ground up by the company’s R&D team and showcasing our commitment to sustainability. The data center is home to ISEG, the most powerful commercially available supercomputer in Europe and the 16th most powerful globally (Top 500 list, November 2023).   

Nebius’s headquarters are in Amsterdam, Netherlands, with teams working out of R&D hubs across Europe and the Middle East.  

Nebius AI is built with the talent of more than 500 highly skilled engineers with a proven track record in developing sophisticated cloud and ML solutions and designing cutting-edge hardware. This allows all the layers of the Nebius AI cloud – from hardware to UI – to be built in-house, distictly differentiating Nebius AI from the majority of specialized clouds: Nebius customers get a true hyperscaler-cloud experience tailored for AI practitioners.

The role

We are seeking a highly skilled ML/AI Engineer to join our team to lead and support benchmarking of GPU platforms for machine learning and AI workloads. You will play a critical role in evaluating the performance of GPU-based hardware for various deep learning and AI frameworks, enabling data-driven decisions for platform optimisation and next-generation hardware development.

 

Your responsibilities will include: 

  • Work closely with hardware, development teams to profile and analyse GPU performance at the system and kernel level.

  • Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g., CUDA, ROCm).

  • Debug and optimise ML workloads to run efficiently on GPU hardware, identifying and resolving performance bottlenecks.

  • Perform acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads.

  • Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system-level optimisations on performance and scalability.

  • Develop tools and dashboards to visualise performance metrics, bottlenecks, and trends.

  • Contribute to internal tooling, frameworks, and best practices

 

We expect you to have:

  • A profound understanding of theoretical foundations of machine learning

  • Deep understanding of performance aspects of large neural networks training and inference (data/tensor/context/expert parallelism, offloading, custom kernels, hardware features, attention optimisations, dynamic batching etc.)

  • Deep experience with modern deep learning frameworks (PyTorch, JAX, Megatron-LM, Tensort-LLM)

  • Good understanding of the GPU stack: CUDA,NCCL, drivers, and relevant libraries

  • Familiarity with containerized environments (e.g., Docker, Kubernetes).

  • Strong communication and ability to work independently

 

Ways to stand out from the crowd:

  • Familiarity with modern LLM inference frameworks (vLLM, SGLang, TensorRT)

  • Experience in Python and performance profiling tools (e.g., Nsight, nvprof, perf).

  • Familiarity with cloud ML platforms like AWS, GCP, Azure ML

  • Contributions to open-source ML benchmarking tools

 
We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us! 
 

 

What we offer 

  • Competitive salary and comprehensive benefits package.
  • Opportunities for professional growth within Nebius.
  • Flexible working arrangements.
  • A dynamic and collaborative work environment that values initiative and innovation.

We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!

Top Skills

Cuda
Docker
Jax
Kubernetes
Nccl
PyTorch
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
473 Employees

What We Do

Cloud platform specifically designed to train AI models

Similar Jobs

Mistral AI Logo Mistral AI

Machine Learning Engineer

Artificial Intelligence
In-Office
5 Locations

Workato Logo Workato

Infrastructure Engineer

Cloud • Enterprise Web • Information Technology • Productivity • Software
In-Office
Amsterdam, NLD
In-Office
6 Locations
In-Office
Amsterdam, NLD

Similar Companies Hiring

Amplify Platform Thumbnail
Fintech • Financial Services • Consulting • Cloud • Business Intelligence • Big Data Analytics
Scottsdale, AZ
62 Employees
Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Standard Template Labs Thumbnail
Software • Information Technology • Artificial Intelligence
New York, NY
10 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account