Accelerator Compiler and Tool Chain Lead

Posted 2 Days Ago
Be an Early Applicant
Santa Clara, CA, USA
In-Office
200K-300K Annually
Entry level
Artificial Intelligence • Hardware • Software • Semiconductor
The Role
Lead architecture and development of an AI accelerator compiler stack, including model ingestion, graph lowering, optimization, quantization, code generation, partitioning, diagnostics, and verification. Partner with hardware, firmware, runtime, and model teams to optimize executable artifacts for NPUs. Lead hiring, mentoring, and technical direction for compiler and ML systems engineers.
Summary Generated by Built In
Role Overview
We are looking for an Accelerator Compiler Lead to own the compiler and model-lowering stack for Velaura’s AI accelerator.This role will lead the path from customer AI models to optimized executable artifacts for our NPU, including graph import, operator lowering, compiler IR, graph transformations, quantization integration, code generation, graph partitioning, and compiler diagnostics. The ideal candidate has built or shipped compiler infrastructure for ML accelerators, GPUs, DSPs, or other heterogeneous compute targets.

    Responsibilities

    ● Lead architecture and development of the AI accelerator compiler stack.

    ● Own model ingestion and graph lowering from frameworks and exchange formats such as PyTorch export flows, ONNX, TensorFlow Lite, or similar.

    ● Define operator coverage strategy, lowering rules, graph transformations, fusion, partitioning, and fallback behavior.

    ● Develop compiler optimization passes for tensor layout, tiling, memory movement, mixed precision, operator fusion, and hardware-specific scheduling.

    ● Work closely with accelerator runtime and driver teams to define executable artifact formats, metadata, memory planning requirements, profiling hooks, and runtime constraints.

    ● Partner with hardware architecture and NPU firmware teams on ISA, command streams, tensor layouts, data movement, hardware constraints, and compiler-visible performance features.

    ● Own quantization compiler integration, including calibration metadata, precision selection, scale handling, layout constraints, and accuracy/performance tradeoffs.

    ● Build compiler diagnostics that help customers understand unsupported operators, shape constraints, graph rewrites, quantization issues, and performance bottlenecks.

    ● Establish compiler verification and regression strategy for graph transformations,IR lowering, numerical behavior, model accuracy, and performance.

    ● Hire, mentor, and lead a team of compiler and ML systems engineers.

    Required Qualifications

    ● Deep experience with compiler development, ML graph compilers, or code generation for accelerators, GPUs, DSPs, or heterogeneous compute systems.

    ● Strong understanding of ML model formats, graph IRs, operator lowering, tensor layouts, quantization, and runtime/compiler interfaces.

    ● Strong C++ and Python programming skills and experience building production-quality compiler or systems software.

    ● Experience with compiler frameworks or technologies such as MLIR, LLVM, TVM, XLA, IREE, Glow, TensorRT-like systems, OpenVINO-like systems, or equivalent.

    ● Strong understanding of correctness risks in compiler optimizations, graph rewrites, mixed precision, operator fusion, and hardware-specific lowering.

    ● Ability to work closely with hardware architects, firmware engineers, runtime engineers, model-integration teams, and SQA.

    ● Experience leading technical teams or major architecture areas.

    Preferred Qualifications

    ● Experience with NPU, GPU, DSP, or AI accelerator compiler stacks.

    ● Experience with quantization-aware compilation, mixed precision, sparsity, pruning, graph partitioning, or hardware-specific scheduling.

    ● Experience supporting ONNX, PyTorch export, TensorFlow Lite, JAX/XLA,TorchDynamo/TorchInductor, or other model import flows.

    ● Familiarity with robotics, computer vision, CNNs, transformers, detection, segmentation, depth, SLAM-adjacent perception, or edge AI workloads.

    ● Experience building customer-facing compiler diagnostics and model-porting tools.

    ● Experience with model-zoo release processes, accuracy validation, and reproducible benchmark artifacts.

    ● Open-source compiler contributions or experience working with external framework communities.

Why Velaura?

Velaura is building next-generation compute technology for cloud, edge, and Physical AI. Our solutions will enable robots, autonomous systems, drones, and other intelligent machines to operate efficiently in the physical world.
This is an opportunity to help build foundational technology at a time when the industry is undergoing fundamental change. You will work alongside experienced leaders, architects, engineers, and operators who have delivered industry-defining products across mobile, cloud, semiconductor, and AI platforms. If you enjoy solving difficult problems, working across disciplines, and helping shape the future of Physical AI, we would love to hear from you.

Equal Employment Opportunity and Accommodations

Velaura is an Equal Opportunity Employer that is committed to inclusion and diversity. Qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, gender, sexual orientation, gender identity, disability or protected veteran status. We also take affirmative action to offer employment
opportunities to minorities, women, individuals with disabilities, and protected veterans.

Velaura is committed to working with qualified individuals with physical or mental disabilities. Applicants who would like to contact us regarding the accessibility of our website or who need special assistance or a reasonable accommodation for any part of the application or hiring process may contact us at: [email protected]. This contact
information is for accommodation requests only. Evaluation of requests for reasonable accommodation will be determined on a case-by-case basis.

Skills Required

  • Deep experience with compiler development, ML graph compilers, or code generation for accelerators, GPUs, DSPs, or heterogeneous compute systems
  • Strong understanding of ML model formats, graph IRs, operator lowering, tensor layouts, quantization, and runtime/compiler interfaces
  • Strong C++ and Python programming skills
  • Experience building production-quality compiler or systems software
  • Experience with compiler frameworks or technologies such as MLIR, LLVM, TVM, XLA, IREE, Glow, TensorRT-like systems, or OpenVINO-like systems
  • Strong understanding of correctness risks in compiler optimizations, graph rewrites, mixed precision, operator fusion, and hardware-specific lowering
  • Ability to work closely with hardware architects, firmware engineers, runtime engineers, model-integration teams, and SQA
  • Experience leading technical teams or major architecture areas
  • Experience with NPU, GPU, DSP, or AI accelerator compiler stacks
  • Experience with quantization-aware compilation, mixed precision, sparsity, pruning, graph partitioning, or hardware-specific scheduling
  • Experience supporting ONNX, PyTorch export, TensorFlow Lite, JAX/XLA, TorchDynamo/TorchInductor, or other model import flows
  • Familiarity with robotics, computer vision, CNNs, transformers, detection, segmentation, depth, SLAM-adjacent perception, or edge AI workloads
  • Experience building customer-facing compiler diagnostics and model-porting tools
  • Experience with model-zoo release processes, accuracy validation, and reproducible benchmark artifacts
  • Open-source compiler contributions or experience working with external framework communities
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
104 Employees
Year Founded: 2022

What We Do

Velaura AI develops ultra-low-power silicon and software for AI compute infrastructure, serving cloud, edge, and physical-AI applications. Its patented, energy-efficient digital-design technology and power-optimized architectures support next-generation AI accelerators, helping compute systems scale with lower energy use. Drawing on deep semiconductor expertise, the company aims to make AI computing more efficient across data centers and intelligent devices, with performance and sustainability central to its approach.

Similar Jobs

Shield AI Logo Shield AI

Engineer I, Software Integration (R6214)

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software • Defense Technology
In-Office
2 Locations
79K-110K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Underwriting Development Program - Global Risk Solutions - June 2027 Start

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Hybrid
11 Locations
40000 Employees
48K-137K Annually

Samsara Logo Samsara

Operations Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
126K-203K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Milpitas, CA, USA
16000 Employees
15-24 Hourly

Similar Companies Hiring

Unusual Machines, Inc. Thumbnail
Hardware • Robotics
Orlando, FL
190 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
65 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account