Senior ML Infrastructure Engineer

Posted 13 Days Ago
Be an Early Applicant
Oxford, Oxfordshire, England, GBR
In-Office
Senior level
Artificial Intelligence • Healthtech • Social Impact • Biotech
The Role
Build and operate high-performance GPU clusters for machine learning training and inference. Design high-throughput compute and storage pathways, optimize I/O and data locality, benchmark distributed systems, and resolve performance bottlenecks. Establish observability, resilience, security automation, quotas, and capacity planning while partnering with research and data teams to scale ML experimentation pipelines.
Summary Generated by Built In

Join us at EIT:

At the Ellison Institute of Technology (EIT), we’re on a mission to translate scientific discovery into real world impact. We bring together visionary scientists, technologists, policy makers, and entrepreneurs to tackle humanity’s greatest challenges in four transformative areas:

  • Health, Medical Science & Generative Biology
  • Food Security & Sustainable Agriculture
  • Climate Change & Managing CO₂
  • Artificial Intelligence & Robotics

This is ambitious work - work that demands curiosity, courage, and a relentless drive to make a difference. At EIT, you’ll join a community built on excellence, innovation, tenacity, trust, and collaboration, where bold ideas become real-world breakthroughs. Together, we push boundaries, embrace complexity, and create solutions to scale ideas for lab to society. Explore more at www.eit.org

Your Role:

Join our SciComp team to build the cloud and compute foundation that enables scientific breakthroughs. Deliver reliable, secure platforms and self-service guardrails that accelerate experimentation and turn ideas into results - faster, at scale, and with confidence. 

Your Responsibilities:

  • Build, operate, and continuously optimise our high-performance GPU training and inference clusters, focusing on robust, high-availability scheduling, isolation, and automated lifecycle management. 
  • Drive systems design and implementation for high-throughput data paths, optimising I/O, caching, and data locality across compute and storage (including our current Lustre implementation). 
  • Proactively benchmark, profile, and resolve performance bottlenecks across the compute, network, and orchestration layers to maximise efficiency for distributed training and inference. 
  • Establish comprehensive observability, resilience, and automated security controls to ensure compliance and robust operation of sensitive research environments. 
  • Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs, setting quotas and streamlining ML experimentation pipelines. 

Requirements

Essential Skills, Qualifications & Experience:

  • Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale 
  • A proactive, autonomous approach to systems design and the proven ability and desire to ideate, co-create and implement optimal solutions 
  • Exposure to migrating or transforming ML infrastructure from traditional schedulers to modern, containerised systems 
  • Expertise with high-throughput storage systems for ML/HPC workloads 
  • Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling to resolve bottlenecks 
  • A solid grasp of IaC and CI/CD practices (e.g., Terraform, Argo CD)

Benefits

We offer the following salary and benefits:

  • Competitive salary (dependent on experience) + travel allowance + bonus
  • Enhanced holiday. Our annual leave allowance is 25 days plus 8 bank holidays and an additional 3 days between Christmas and New Year. You will also have the opportunity to purchase an additional 5 days annual leave in January and July.
  • Pension - Employer contribution 7.5%, minimum employee contribution 5%
  • Life Assurance.
  • Income Protection
  • Private Medical Insurance as standard for you, your partner and any dependents. Including hospital Cash Plan
  • Employee discounts
  • Electric car scheme
  • Nursery Salary Sacrifice scheme
  • Cycle to Work Scheme
  • Family Planning
  • Neurodiversity support including advise and assessments
  • Coaching & Therapy services

Working together - what it involves:

You must have the right to work permanently in the UK with a willingness to travel as necessary. In certain cases, we can consider sponsorship, and this will be assessed on a case-by-case basis.

You will live in, or within easy commuting distance of, Oxford/London (or be willing to relocate)

The SciComp team work to a hybrid working pattern of 3 days in the office, 2x in our Oxford Office and 1x in our London office. You must be able to commit to this should you be successfully appointed.

Skills Required

  • Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale
  • Proactive, autonomous approach to systems design, with the ability to ideate, collaborate, and implement optimal solutions
  • Experience migrating or transforming ML infrastructure from traditional schedulers to modern containerized systems
  • Expertise with high-throughput storage systems for ML or HPC workloads
  • Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling
  • Understanding of Infrastructure as Code and CI/CD practices, such as Terraform and Argo CD
  • Right to work permanently in the UK
  • Residence within commuting distance of Oxford or willingness to relocate
  • Ability to work onsite at the Oxford office at least three days per working week
  • Willingness to travel as necessary
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Oxford
Year Founded: 2023

What We Do

The Ellison Institute of Technology aims to discover, develop, and deploy science and technology to solve humanity's most important problems, focusing on health and medical science, food security, climate change, and AI-driven government innovation.

Similar Jobs

Roku Logo Roku

Senior Software Engineer

News + Entertainment
In-Office
Cambridge, Cambridgeshire, England, GBR
2724 Employees

Morningstar Logo Morningstar

Senior Product Marketing Manager

Artificial Intelligence • Big Data • Enterprise Web • Fintech • Software • Financial Services
Hybrid
2 Locations
11500 Employees
121K-186K Annually

Ericsson Logo Ericsson

Consultant

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Reading, Berkshire, England, GBR
88000 Employees

Boeing Logo Boeing

Support Engineer

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office
Waddington, Ribble Valley, Lancashire, England, GBR
170000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account