Senior MLOps Engineer

Reposted One Month Ago
Palo Alto, CA, USA
In-Office
Senior level
Artificial Intelligence • Enterprise Web • Software • Generative AI
The Role
Own and operate end-to-end ML infrastructure for training, serving, and evaluation of LLM/SLM models. Build scalable low-latency inference (vLLM, batching, autoscaling), multi-GPU training/serving clusters, observability and monitoring, inference-time optimizations (quantization, distillation), reproducibility, versioning, and enterprise-grade auditability. Set MLOps best practices and standards.
Summary Generated by Built In

Palo Alto, CA | Full-Time | On-site

About Us:

At Nace AI, we are redefining how professional services operate by delivering Sovereign Specialized Intelligence. As an applied research and product company, we equip enterprises with a comprehensive AI stack to build customized, secure intelligence tailored to their unique business needs.

Driven by advanced Small Language Models and our dynamic metamodel framework, our flagship platforms - Nace Data Intelligence and the Nace SLM Cloud - enable true end-to-end business process automation. The result is transformative ROI: professional services firms using Nace AI are currently recovering 1,000 hours per client engagement, drastically reducing overhead and accelerating delivery.

The work we are doing has a meaningful impact across industries, and every hire at Nace AI plays a critical role in shaping the company’s trajectory. This is a unique opportunity to join a high conviction AI company at an early stage and directly influence its growth.

 

If building a world-class AI team from the ground up excites you, we’d love to talk.

Role Overview:

As a Senior MLOps Engineer, you will own the infrastructure that takes Nace.AI's models from research to reliable, production-grade systems. Our infrastructure generates task-specific Small Language Models (SLMs) in real time — which means our training, serving, and evaluation infrastructure isn't an afterthought; it is the product. You will design and operate the pipelines, orchestration, and serving layers that allow us to train, deploy, monitor, and continuously improve many specialized models at once, with the reliability that high-stakes audit, compliance, and finance workflows demand. This role sits at the intersection of ML engineering, LLM inference infrastructure, and platform reliability, and requires both strong systems instincts and hands-on execution.

Key Responsibilities:

  • Design, build, and operate end-to-end ML infrastructure: training orchestration, experiment tracking, model registries, CI/CD for models, and automated evaluation pipelines.

  • Own LLM/SLM serving infrastructure — scale low-latency, high-throughput inference using frameworks like vLLM, including batching, caching, and autoscaling strategies.

  • Build and manage multi-GPU training and inference clusters (scheduling, utilization, cost optimization) across cloud and on-prem environments.

  • Implement observability for models in production: latency, throughput, drift, regression, and quality monitoring with actionable alerting.

  • Apply inference-time optimizations — quantization (AWQ, GPTQ, FP8/GGUF), distillation support, KV-cache management, and deployment tuning — in partnership with our ML and Research Engineers.

  • Harden our stack for enterprise deployment: reproducibility, versioning, access controls, and audit-ready traceability of model behavior.

  • Set MLOps best practices and tooling standards as an early, senior member of the infrastructure team.

Qualifications:

  • 5+ years of experience in MLOps, ML infrastructure, or platform engineering, with substantial production ownership.

  • Proven experience deploying and scaling LLM, inference infrastructure in production, including model serving frameworks such as TRT, vLLM, SGLang or TGI.

  • Strong proficiency with Kubernetes, containerization (Docker), and infrastructure-as-code (Terraform or similar).

  • Hands-on experience with GPU cluster management and distributed training/serving environments.

  • Proficient in Python with a strong track record of building substantial, maintainable systems.

  • Experience with ML pipeline and orchestration tooling (e.g., Airflow, Kubeflow, Ray, MLflow, Weights & Biases).

  • Solid foundation in computer science fundamentals and cloud architecture (AWS, GCP, or Azure).

  • BS degree in CS or related technical field.

  • Self-starter comfortable working in a fast-paced, dynamic environment.

Preferred Qualifications:

  • MS in CS or related technical field.

  • Experience operating multi-node GPU training infrastructure.

  • Hands-on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF) and other inference-time optimizations.

  • Familiarity with data processing stacks such as Spark and Airflow.

  • Experience supporting fine-tuning workflows for LLMs/VLMs (instruction tuning, RLHF/DPO pipelines).

  • Experience in regulated or enterprise environments where reliability, security, and auditability are first-class requirements.

  • Contributor to open-source ML infrastructure projects.

Why Nace AI?

  • Pedigree: Work with a team from top-tier institutions and companies, backed by the best VCs in the world.

  • Impact: You are joining early enough to shape the infrastructure foundations of a company aiming to be the "OS" for professional knowledge.

  • Competitive Package: Silicon Valley-standard salary, significant equity, and premium benefits.

Skills Required

  • 5+ years experience in MLOps, ML infrastructure, or platform engineering with production ownership
  • Proven experience deploying and scaling LLM inference infrastructure (e.g., TRT, vLLM, SGLang, TGI)
  • Strong proficiency with Kubernetes and containerization (Docker)
  • Experience with infrastructure-as-code (Terraform or similar)
  • Hands-on experience with GPU cluster management and distributed training/serving
  • Proficient in Python and building maintainable systems
  • Experience with ML pipeline and orchestration tooling (Airflow, Kubeflow, Ray, MLflow, Weights & Biases)
  • Solid foundation in computer science fundamentals and cloud architecture (AWS, GCP, or Azure)
  • BS degree in Computer Science or related technical field
  • Self-starter comfortable working in a fast-paced, dynamic environment
  • MS in Computer Science or related technical field
  • Experience operating multi-node GPU training infrastructure
  • Hands-on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF) and inference-time optimizations
  • Familiarity with data processing stacks such as Spark
  • Experience supporting fine-tuning workflows for LLMs/VLMs (instruction tuning, RLHF/DPO)
  • Experience in regulated or enterprise environments emphasizing reliability, security, and auditability
  • Contributor to open-source ML infrastructure projects
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
50 Employees
Year Founded: 2024

What We Do

Nace.AI develops AI systems that generate custom, task-specific AI models for enterprises, focusing on applications in audit, compliance, and professional deliverables.

Similar Jobs

In-Office or Remote
2 Locations
300 Employees
152K-230K Annually

NVIDIA Logo NVIDIA

Senior MLOps Engineer - DSX Enablement

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
2 Locations
21960 Employees
184K-357K Annually

Intuitive Logo Intuitive

Sr MLOps Engineer

Healthtech • Robotics
In-Office
Sunnyvale, CA, USA
12000 Employees

ClimateAI Logo ClimateAI

Senior MLOps Engineer

Artificial Intelligence • Agriculture
In-Office
San Francisco, CA, USA
64 Employees
170K-200K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account