ML Engineer (Senior/Staff)

Posted Yesterday
Hiring Remotely in Seattle, WA, USA
In-Office or Remote
140K-220K Annually
Senior level
Artificial Intelligence • Industrial
The Role
Own the architecture and technical direction for production ML inference and evaluation infrastructure. Responsibilities include GPU scheduling, Kubernetes-based autoscaling, model serving with vLLM and SGLang, cost-aware model routing, LLM evaluation pipelines, golden datasets, regression gates, and reliability engineering. The role also applies physics-informed ML to client and platform problems, establishes organizational standards, mentors engineers, and partners across infrastructure teams. This is a fully remote senior/staff individual contributor role limited to candidates in the United States and Canada.
Summary Generated by Built In

About AZX

Our mission is to accelerate positive impact in critical industries through AI transformation. We specialize in physics-informed ML and enterprise AI solutions that directly address climate and sustainability challenges.

We’re growing quickly and already work with category-leaders in real estate (CBRE), energy (LevelTen Energy), logistics (Flexe) and utilities.

We’re a public benefit corporation, founded in 2024, and have been profitable from inception.

We work on challenges in clean energy, decarbonization, climate risk, energy systems, and global economics. We’re building our company for long-term success and aim to create the ultimate place to work for those passionate about AI and making a positive impact.

About the Role

We're looking for an ML Engineer to own the technical backbone of how AZX serves and evaluates models at scale. This is a high-leverage IC role spanning our inference platform — GPU scheduling, autoscaling, and serving infrastructure for vLLM/SGLang across cloud and customer-managed clusters — and the evaluation systems that tell us whether model, prompt, and agent changes actually make things better.

You'll create technical direction for how AZX serves models reliably. This role suits someone who wants architectural ownership over hard ML infrastructure problems, paired with the judgment to build the guardrails that let the rest of the team move fast safely.

Responsibilities:

  • Own architecture for inference serving and GPU scheduling — Kubernetes operators, autoscaling, and dynamic capacity across vLLM/SGLang deployments on cloud and customer-managed infrastructure.

  • Design and calibrate eval systems for model, prompt, and agent changes, including golden datasets, LLM-as-judge pipelines, and regression gates wired into CI.

  • Advise on cost-aware model routing and cascading decisions, balancing latency, cost, and quality across providers and model tiers.

  • Apply physics-informed ML and enterprise AI expertise to the hardest client and platform problems, drawing on the team's research depth.

  • Set technical standards for ML infrastructure and evaluation practice across the org, and mentor engineers working in this space.

  • Partner closely with the inference platform, gateway, and evals-focused engineers to keep architecture coherent as the platform grows.

Core Qualifications:

  • 3+ years of experience with ML infrastructure and inference serving — vLLM, SGLang, TensorRT-LLM, or comparable systems — at production scale.

  • Strong background in evaluation and reliability engineering for ML/LLM systems, or the seniority to build this practice from scratch.

  • Solid Kubernetes experience, ideally including GPU-specific scheduling constraints (node pools, autoscaling under GPU bottlenecks).

  • A track record of technical leadership at a staff or senior level — setting direction, not just executing tickets.

  • Research fluency is a plus (PhD, publications, or equivalent depth) given the technical bar of our existing ML team, though this is an infrastructure-and-systems role first.

Bonus Qualifications:

  • Advanced ML/AI frameworks and techniques (e.g., PyTorch Lightning, JAX, HuggingFace, ONNX optimizations)

  • Lower-level or performance-focused languages for ML acceleration (e.g., C++, Rust, CUDA)

  • Large-scale data and distributed training paradigms (e.g., Spark, Ray, Horovod, Dask)

  • Advanced data infrastructure (e.g., vector/graph databases, feature stores, data lakes)

Why AZX!

  • Be part of a fast-growing, profitable, mission-driven company with industry-leading clients tackling the massive opportunity of AI transformation in critical industries.

  • Competitive early-stage startup compensation (based on capabilities, experience, and location)

  • Bonus eligibility

  • Health insurance with meaningful coverage for dependents

  • Flexible paid time off

  • Equity

  • Fully remote culture with a cluster of teammates in Seattle

Additional Information:

  • Must be willing to travel to Seattle area for final interview and travel 2x/year for company summits

  • Applicants must be currently authorized to work in the United States on a full-time basis.

  • We are unable to sponsor or take over sponsorship of employment visas at this time.

Next Steps:

If this job sounds like a great fit but don’t check ALL of these qualification boxes, we’d still love to hear from you!

Skills Required

  • 3+ years of experience with ML infrastructure and production-scale inference serving, including vLLM, SGLang, TensorRT-LLM, or comparable systems
  • Strong background in evaluation and reliability engineering for ML or LLM systems, or equivalent seniority to build the practice from scratch
  • Solid Kubernetes experience, ideally including GPU scheduling, node pools, and autoscaling under GPU bottlenecks
  • Technical leadership experience at a senior or staff level, including setting technical direction
  • High emotional intelligence and a learning mindset
  • Strong collaboration skills
  • Comfort making decisions amid ambiguity and course correcting
  • Research fluency through a PhD, publications, or equivalent depth
  • Experience in both startup and enterprise environments
  • Experience in energy, real estate, utilities, climate, or related fields
  • Experience with advanced ML frameworks and techniques such as PyTorch Lightning, JAX, Hugging Face, or ONNX optimizations
  • Experience with C++, Rust, or CUDA for ML acceleration
  • Experience with Spark, Ray, Horovod, or Dask for large-scale data and distributed training
  • Experience with vector or graph databases, feature stores, or data lakes
  • Current authorization to work full-time in the United States
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Seattle, Wa
9 Employees
Year Founded: 2024

What We Do

Tech transformation, societal shifts and environmental disruption put enormous pressure on critical industries like energy, infrastructure, real estate and others to adapt, pursue new growth opportunities and optimize operations. AZX was created to meet these challenges head-on. We deliver AI transformation for critical industries providing strategy, execution and acceleration solutions. We prioritize code over powerpoint, applying start-up principles to challenge conventional thinking and develop strategic AI technology and business solutions that can be implemented quickly. We are mission-driven and client-centric, solving hard problems and building value for people, organizations and the planet. We are AZX.

Similar Jobs

Airbnb Logo Airbnb

Machine Learning Engineer

Real Estate • Travel • PropTech
Remote
USA
14622 Employees
248K-310K Annually

Airbnb Logo Airbnb

Machine Learning Engineer

Real Estate • Travel • PropTech
Remote
USA
14622 Employees
244K-305K Annually

Airbnb Logo Airbnb

Machine Learning Engineer

Real Estate • Travel • PropTech
Remote
United States
14622 Employees
244K-305K Annually

Nextdoor Logo Nextdoor

Machine Learning Engineer

Information Technology • Other • Social Media
Remote or Hybrid
US
780 Employees
190K-355K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account