ML Ops Engineer

Posted 4 Days Ago
Be an Early Applicant
London, Greater London, England, GBR
In-Office
Senior level
Information Technology
The Role
Design, scale, and maintain cloud-native MLOps and LLMOps infrastructure. Provision GPU orchestration, optimise compute/network/storage, build CI/CD and MLOps pipelines, deploy LLMs with inference engines, automate with IaC, monitor model/data drift and latency, and manage cloud GPU cost and observability.
Summary Generated by Built In

At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.

What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.

Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.

Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebrating our wins – big and small.

Supported by operating principles of being strategy-led, values-based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!

Role Overview

We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused scenario planning platform.

You will work closely with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while ensuring optimal GPU utilisation, reliability, and cost-efficiency.

Your Impact

  • Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).
  • Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible.
  • Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference.
  • Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment.
  • Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).
  • Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.
  • Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure.
  • Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste.
  • Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models.
  • Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow.

Your Skills

  • Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure.
  • Proven track record of deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production environments.
  • Demonstrated experience managing compute-intensive GPU infrastructure and high-performance computing (HPC) environments.
  • Advanced proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).
  • Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
  • Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI.
  • Solid background in AWS / GCP / Azure, Kubecost, and GPU cost optimisation techniques.
  • Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance monitoring.

Our Commitment to Diversity, Equity, Inclusion and Belonging (DEIB)

We believe attracting and retaining the best talent and fostering an inclusive culture strengthens our business. DEIB improves our workforce, enhances trust with our partners and customers, and drives business success. Build your career in a place where diversity, equity, inclusion and belonging aren’t just words on paper – this is what drives our innovation, it’s how we connect, and it contributes to what makes us a market leader. We believe in a hiring and working environment where all people are respected and valued, regardless of gender identity or expression, sexual orientation, religion, ethnicity, age, neurodiversity, disability status, citizenship, or any other aspect which makes people unique. We hire you for who you are, and we want you to bring your authentic self to work every day! 

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive equitable benefits and all privileges of employment. Please contact us to request accommodation.  

Fraud Recruitment Disclaimer  

It has come to our attention that fraudulent and fictitious job opportunities are being circulated on the Internet. Prospective candidates are being contacted by certain individuals, mainly through telephone calls, emails and correspondence, claiming they are representatives of Anaplan. The main purpose of these correspondences and announcements is to obtain privileged information from individuals.  

Anaplan does not:  

  • Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person.   
  • Send job offers via email. All offers are first extended verbally by a member of our internal recruitment team whenever possible and then followed up via written communication.  

All emails from Anaplan would come from an @anaplan.com email address. Should you have any doubts about the authenticity of an email, letter or telephone communication purportedly from, for, or on behalf of Anaplan, please send an email to [email protected] before taking any further action in relation to the correspondence.   

Candidate data processed during our recruitment activities is handled in accordance with our Candidate Privacy Notice. This may include the use of artificial intelligence or automated tools to assist our team in evaluating qualifications.


Skills Required

  • Production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering with AI/ML infrastructure exposure
  • Proven track record deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production
  • Experience managing compute-intensive GPU infrastructure and high-performance computing (HPC) environments
  • Advanced proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio)
  • Hands-on experience with Infrastructure as Code and automation tools such as Terraform and Ansible
  • Experience with CI/CD tools (GitHub Actions, ArgoCD, or Jenkins) and building MLOps pipelines
  • Experience with LLM and inference tooling (Triton Inference Server, vLLM, TensorRT-LLM, DeepSpeed, Hugging Face TGI)
  • Familiarity with ML orchestration and tooling (Ray, MLflow, LangChain / LangSmith, Weights & Biases)
  • Solid background administering cloud platforms (AWS, GCP, or Azure) and tools for cloud cost optimisation (Kubecost)
  • Strong programming/scripting skills in Python, Bash, or Go
  • Deep knowledge of Linux kernel tuning and performance monitoring
  • Experience optimising GPU workloads, high-speed networking, and storage for model training and low-latency inference
  • Experience implementing observability and monitoring using Prometheus, Grafana, OpenTelemetry, and model monitoring tools
  • Experience implementing auto-scaling strategies, spot instance policies, and dynamic resource allocation to control cloud spend

Anaplan Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Anaplan and has not been reviewed or approved by Anaplan.

  • Strong & Reliable Incentives Earnings potential in sales is considered good or fair, with on‑target earnings achievable when plans are met. This indicates variable pay can meaningfully boost total compensation when targets are hit.
  • Healthcare Strength Medical, dental, and vision coverage are described as strong, complemented by mental‑health resources and EAP support. Company‑wide paid wellbeing days reinforce the health and wellness focus.
  • Parental & Family Support Equitable parental and caregiver leave are highlighted alongside fertility and adoption support. Family‑forming programs such as Carrot are part of the offering.

Anaplan Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Miami, FL
2,194 Employees
Year Founded: 2006

What We Do

Anaplan is building a future where connected leaders and teams are able to constantly adapt, transform and reinvent their businesses. We make it possible to share actionable insights, empower and unleash creativity, and drive innovation. With Anaplan, finance and operational leaders across the organization can model complex scenarios, forecast continuously with added intelligence, and make agile decisions with confidence.

Similar Jobs

Humanoid Logo Humanoid

Senior Dev/ML Ops Engineer

Artificial Intelligence • Robotics
In-Office
London, Greater London, England, GBR
200 Employees

Preply Logo Preply

Staff Machine Learning Ops Engineer

Edtech • Information Technology • Software
Hybrid
London, Greater London, England, GBR
700 Employees

Circadia Health Logo Circadia Health

ML Ops Engineer

Artificial Intelligence • Healthtech
In-Office
London, Greater London, England, GBR
120 Employees
Hybrid
2 Locations
126 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account