Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

Posted 3 Days Ago
Be an Early Applicant
San Jose, CA, USA
In-Office
170K-339K Annually
Senior level
Artificial Intelligence • Mobile • Other • Transportation
The Role
Lead deployment, optimization, and resource scheduling of vehicle and cloud AI models. Design low-latency inference pipelines, build system stability and profiling frameworks, optimize GPU execution, and scale LLM serving for simulation and validation. Collaborate with perception, cloud, and safety teams to enable reliable, production-grade autonomous deployment.
Summary Generated by Built In

About the Company

DiDi's autonomous driving unit was established in 2016 with the mission of developing Level 4 autonomous driving (AD) technology to make transportation safer and more efficient. In August 2019, the unit became an independent company, DiDi Autonomous Driving, dedicated to advanced AD R&D, product application, and business expansion. We believe integrating AD technology into a shared-mobility fleet will generate immense social value. By leveraging DiDi's specialized technology, operational expertise, and integrated ecosystem, we are positioned to build and operate a highly efficient, user-oriented autonomous fleet.

 

About The Role

We are seeking an experienced and mission-driven Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization to lead the performance tuning, deployment, and resource scheduling of cutting-edge AI models across on-vehicle and cloud infrastructure. In this role, you will design high-efficiency inference pipelines, build system-level stability frameworks, and optimize hardware execution to ensure ultra-low latency and rock-solid operational reliability. You will act as a technical leader in AI infrastructure, accelerating model iteration and bridging the gap between frontier deep learning algorithms and real-time autonomous systems.

 

Responsibilities

  • Own the deployment, optimization, and resource scheduling of vehicle-side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.

  • Lead vehicle-side system stability initiatives, conducting independent root-cause analysis and driving resolution for complex, system-level performance bottlenecks and runtime anomalies.

  • Architect and scale service-oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.

  • Track and evaluate cutting-edge industry methodologies, continuously integrating advanced optimization toolchains, quantization techniques, and execution engines.

  • Establish system-level profiling and telemetry frameworks using CUDA tools to monitor, analyze, and maximize hardware utilization across target GPU architectures.

  • Collaborate cross-functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

 

Qualifications

  • Master’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.

  • 3-8+ years of industry experience in high-performance computing, AI infrastructure, model optimization, or embedded deployment.

  • Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools.

  • Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM).

  • Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.

  • Demonstrated ability to diagnose complex software-hardware integration issues and drive scalable, production-grade solutions.

Preferred Qualifications

  • Hands-on experience optimizing and deploying AI models on the NVIDIA Thor platform, including hardware resource scheduling and acceleration.

  • Proven track record of serving large foundation models (e.g., LLaMA, Qwen, GPT) in production or high-throughput cloud pipelines using frameworks like vLLM, SGLang, TGI, or LightLLM.

  • Background in deep learning training frameworks (PyTorch) and practical experience with model quantization (INT8/FP8/AWQ), kernel fusion, or graph compilation.

  • Experience deploying real-time, high-availability AI workloads in autonomous vehicles, robotics, or edge devices.

 

The base salary range for this full-time position is $169,783 - $351,000 annually in addition to bonus, equity and benefits. Our salary ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
I acknowledge that prior to submitting this application, I have read and accepted the Privacy Notice for California Residents which is available on https://v.didi.cn/AQnxlBa

Skills Required

  • Bachelor's or higher degree in Computer Science, Software Engineering, Systems Engineering, or related field
  • 3-8+ years industry experience in high-performance computing, AI infrastructure, model optimization, or embedded deployment
  • Strong proficiency in C++ and Python
  • Expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools
  • Familiarity with inference engines (TensorRT, ONNX Runtime) and LLM inference/serving frameworks (vLLM, SGLang, TensorRT-LLM)
  • Practical understanding of modern GPU architectures (NVIDIA Hopper, Thor) and memory bandwidth management
  • Ability to diagnose software-hardware integration issues and deliver scalable production solutions
  • Hands-on experience optimizing and deploying AI models on the NVIDIA Thor platform, including resource scheduling and acceleration
  • Proven experience serving large foundation models (LLaMA, Qwen, GPT) in production or high-throughput pipelines using vLLM, SGLang, TGI, or LightLLM
  • Background with PyTorch, model quantization (INT8/FP8/AWQ), kernel fusion, or graph compilation
  • Experience deploying real-time, high-availability AI workloads in autonomous vehicles, robotics, or edge devices
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Heerhugowaard
16,230 Employees
Year Founded: 2012

What We Do

DiDi Global Inc. (NYSE: DIDI) is the world’s leading mobility technology platform. It offers a wide range of app-based services across Asia-Pacific, Latin America and Africa, as well as in Central Asia and Russia, including ride hailing, taxi hailing, chauffeur, hitch and other forms of shared mobility as well as auto solutions, food delivery, intra-city freight and financial services. DiDi provides car owners, drivers and delivery partners with flexible work and income opportunities. It is committed to collaborating with policymakers, the taxi industry, the automobile industry and the communities to solve the world’s transportation, environmental and employment challenges through the use of AI technology and localized smart transportation innovations. DiDi strives to create better life experiences and greater social value, by building a safe, inclusive and sustainable transportation and local services ecosystem for cities of the future.

Similar Jobs

Coursera + Udemy  Logo Coursera + Udemy

Principal Product Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
United States
1500 Employees
219K-274K Annually

Mondelēz International Logo Mondelēz International

Product Owner

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
United States
90000 Employees
140K-193K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Associate Claims Adjuster, Workers Compensation

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Hybrid
11 Locations
40000 Employees
50K-94K Annually

Coursera + Udemy  Logo Coursera + Udemy

Senior Category Growth Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
United States
1500 Employees
137K-174K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account