Inference

Posted 2 Months Ago
Be an Early Applicant
2 Locations
Hybrid
Senior level
Artificial Intelligence • Machine Learning • Robotics
The Role
Design and implement low-latency, high-throughput inference pipelines for on-device and GPU-cluster deployment. Optimize kernels, batching, quantization, memory and scheduling. Build distributed serving systems, integrate low-level CUDA/Triton code with high-level frameworks, and develop monitoring/debugging tools to ensure reliability and determinism.
Summary Generated by Built In
What You’ll Do
  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics

  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization

  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks

  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)

  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)

  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)

  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling

  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments

  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

Skills Required

  • 8+ years experience in distributed systems, ML infrastructure, or high-performance serving
  • Production-grade expertise in Python
  • Strong background in systems languages (C++, Rust, or Go)
  • Low-level performance mastery (CUDA, Triton, kernel optimization, custom kernels)
  • Experience with quantization, memory and compute scheduling, and graph compilation
  • Proven track record scaling inference for both throughput-oriented clusters and latency-critical on-device deployments
  • History of tuning hardware-software interactions for efficiency, throughput, and responsiveness
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
86 Employees
Year Founded: 2025

What We Do

Genesis AI is a global full-stack robotics company developing general-purpose robots with human-level intelligence and capabilities. It aims to build foundational AI models that automate repetitive tasks across applications such as lab work and housekeeping. The company uses a proprietary physics engine to generate synthetic physical-world data, helping train robotics models for diverse real-world environments, and operates across Paris and Silicon Valley.

Similar Jobs

Scaleway Logo Scaleway

Product Manager

Cloud • Software
Hybrid
Paris, Île-de-France, FRA
555 Employees

Dragonfly Logo Dragonfly

Senior Inference Optimization Engineer - Dragonfly Portfolio

Fintech • Software • Financial Services • Cryptocurrency
Remote or Hybrid
34 Locations
152 Employees

NVIDIA Logo NVIDIA

Senior Software Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Remote or Hybrid
4 Locations
21960 Employees
293K-650K Annually

Similar Companies Hiring

Unusual Machines, Inc. Thumbnail
Hardware • Robotics
Orlando, FL
190 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
65 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account