Inference Optimization Engineer

Posted Yesterday
San Francisco, CA, USA
In-Office
Entry level
Consumer Web • Digital Media • Enterprise Web • Marketing Tech • News + Entertainment • Software • Generative AI
A generative media company building the AI-native creation platform around the world's first omnimodal foundation model.
The Role
Optimize visual AI model inference across architectures, algorithms, runtimes, GPU kernels, and distributed systems. Profile workloads, identify compute and memory bottlenecks, improve latency, throughput, memory efficiency, and GPU utilization, and develop benchmarking infrastructure. The role involves techniques such as quantization, sparsity, caching, compilation, attention optimization, custom CUDA or Triton kernels, and multi-GPU execution while collaborating closely with research scientists.
Summary Generated by Built In
About Hedra

Hedra is the platform, models, and infrastructure for visual intelligence.

We build models and systems that push the frontier of visual intelligence, along with the infrastructure required to make those models fast, efficient, reliable, and accessible at scale.

We’re a small, highly technical team in San Francisco, backed by a16z and other leading investors. Researchers and engineers at Hedra work closely across boundaries, own problems end to end, and have significant influence over both what we build and how we build it.

The Role

We’re looking for an Inference Optimization Engineer to work alongside our research team on making state-of-the-art visual models fast and efficient at inference time.

You’ll work at the boundary between research and systems, taking new model architectures and figuring out how to run them efficiently on modern hardware. That means understanding where time and memory are being spent, identifying opportunities for algorithmic and systems-level improvements, and implementing optimizations across model architecture, inference algorithms, runtimes, kernels, and distributed execution.

The problems rarely live neatly within one layer of the stack. Depending on what you find, you might modify how a model executes, develop a new inference technique, write a custom GPU kernel, rethink memory movement, or change how work is distributed across accelerators.

We care more about technical depth, curiosity, and demonstrated ability than years of experience. We’re open to experienced ML systems engineers as well as exceptional early-career engineers or researchers who have already gone unusually deep on model performance, GPU systems, or efficient inference.

What You’ll Do
  • Work directly with research scientists and engineers to make new visual models fast and efficient at inference time.

  • Profile model architectures and workloads to understand bottlenecks across compute, memory, communication, and model execution.

  • Develop and implement new approaches to improving inference latency, throughput, memory efficiency, and GPU utilization.

  • Explore algorithmic optimizations including quantization, sparsity, caching, compilation, attention optimizations, and alternative execution strategies.

  • Build or optimize GPU kernels using CUDA, Triton, or similar technologies when existing implementations leave performance on the table.

  • Optimize model execution across single-GPU, multi-GPU, and multi-node environments.

  • Reason about the interaction between model architecture and hardware, and work with researchers when architectural changes can unlock meaningful performance improvements.

  • Investigate communication, memory movement, parallelism, and distributed execution strategies for large visual models.

  • Build rigorous benchmarking, profiling, and performance-regression infrastructure to understand performance and evaluate new optimization ideas.

  • Evaluate new inference runtimes, compilers, frameworks, optimization techniques, and accelerator hardware.

  • Stay close to advances in efficient inference, GPU programming, model architectures, compilers, and ML systems research, and rapidly test promising ideas.

  • Help turn research breakthroughs into models that can be deployed and served efficiently at scale.

What We’re Looking For
  • Deep technical ability in efficient ML inference, ML systems, GPU computing, or adjacent research, demonstrated through research, production engineering, open-source contributions, or unusually ambitious independent work.

  • Strong understanding of how modern deep learning models execute on hardware, including the relationship between compute, memory, communication, and performance.

  • Experience profiling ML workloads, identifying bottlenecks, forming hypotheses, and driving measurable performance improvements.

  • Strong programming fundamentals in Python, C++, or another systems-oriented language.

  • Experience with some combination of PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, or comparable technologies.

  • Ability to reason across abstraction layers rather than treating model architecture, framework, runtime, kernel, and hardware boundaries as fixed.

  • Strong intuition for performance tradeoffs across latency, throughput, memory, numerical precision, model quality, and complexity.

  • Curiosity about how models work internally and a willingness to modify or rethink existing approaches when the performance problem calls for it.

  • Comfort working on ambiguous problems where the bottleneck, and sometimes even the right question, is not known in advance.

  • Ability to communicate technical ideas clearly and collaborate closely with research scientists and engineers.

We don’t expect every candidate to have experience across the entire stack. Exceptional depth in one or more relevant areas, combined with the ability and curiosity to reason across the others, matters more to us than checking every box.

Nice to Have
  • Experience optimizing large generative, multimodal, vision, or video models.

  • CUDA, Triton, CUTLASS, or other GPU kernel development.

  • Deep knowledge of GPU architecture, memory hierarchy, and hardware-aware optimization.

  • Experience with attention optimization, kernel fusion, memory-efficient execution, or custom operators.

  • Model compilation or graph optimization experience.

  • Quantization, sparsity, caching, speculative execution, or other efficient inference techniques.

  • Experience optimizing diffusion, autoregressive, transformer, or other large generative architectures.

  • Multi-GPU or multi-node model execution, including tensor, pipeline, sequence, or other forms of parallelism.

  • Experience optimizing communication or data movement between accelerators.

  • Experience with profiling tools such as Nsight Systems or Nsight Compute.

  • Contributions to ML systems, inference runtimes, compilers, GPU libraries, or performance-focused open-source projects.

  • Research or publications in efficient ML, ML systems, GPU computing, compilers, or related areas.

Benefits:
  • Competitive compensation and equity

  • 401k

  • Healthcare (Silver PPO Medical, Vision, Dental)

  • Lunch and snacks at the office

This role is based in San Francisco, and we work together in person five days a week.

Skills Required

  • Deep technical ability in efficient ML inference, ML systems, GPU computing, or adjacent research
  • Strong understanding of modern deep learning model execution on hardware, including compute, memory, communication, and performance
  • Experience profiling ML workloads, identifying bottlenecks, and driving measurable performance improvements
  • Strong programming fundamentals in Python, C++, or another systems-oriented language
  • Experience with some combination of PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, or comparable technologies
  • Ability to reason across model architecture, framework, runtime, kernel, and hardware abstraction layers
  • Understanding of performance tradeoffs involving latency, throughput, memory, numerical precision, model quality, and complexity
  • Ability to communicate technical ideas clearly and collaborate with research scientists and engineers
  • Experience optimizing large generative, multimodal, vision, or video models
  • CUDA, Triton, CUTLASS, or other GPU kernel development experience
  • Experience with attention optimization, kernel fusion, memory-efficient execution, or custom operators
  • Model compilation, graph optimization, quantization, sparsity, caching, or speculative execution experience
  • Multi-GPU or multi-node model execution and accelerator communication optimization experience
  • Experience with Nsight Systems, Nsight Compute, or similar profiling tools
  • Research or publications in efficient ML, ML systems, GPU computing, compilers, or related fields

Hedra Compensation & Benefits Highlights

  • Healthcare Strength — Medical, dental, and vision insurance are consistently listed, with specifics such as a Silver PPO medical plan noted in public materials. This points to strong core health coverage for a small, early-stage startup.
  • Equity Value & Accessibility — Company equity is included as part of total compensation across roles, aligning with high-upside packages typical of early-stage AI startups. Role descriptions frequently pair competitive base pay with stock options.
  • Leave & Time Off Breadth — An unlimited PTO policy is described across public job and profile listings. This indicates broad time-off flexibility, even if actual usage norms are not detailed.

Hedra Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
14 Employees
Year Founded: 2023

What We Do

Hedra is an AI native platform for multimodal creation. The platform is built around their own cutting-edge proprietary video model, Character-3, which is the first multimodal model in production. Alongside Character-3, the platform also brings other leading foundation models into one ecosystem spanning generative images, video, and audio. Prosumer and enterprise users leverage Hedra to generate content ranging from viral social media to branded content marketing.

Why Work With Us

We're an early-stage team that moves very fast and is building at the leading edge of AI/Media. Every employee takes on a lot of ownership and has an opportunity to learn and grow rapidly.

Gallery

Gallery

Hedra Offices

OnSite Workspace

Hedra's main office is in San Francisco and secondary hub is in New York.

Typical time on-site: None
HQHQ
New York, New York
Learn more

Similar Jobs

Hedra Logo Hedra

Product Marketing Lead

Consumer Web • Digital Media • Enterprise Web • Marketing Tech • News + Entertainment • Software • Generative AI
In-Office
San Francisco, CA, USA
14 Employees

Hedra Logo Hedra

Head Of Marketing

Consumer Web • Digital Media • Enterprise Web • Marketing Tech • News + Entertainment • Software • Generative AI
In-Office
San Francisco, CA, USA
14 Employees

Hedra Logo Hedra

Research Engineer

Consumer Web • Digital Media • Enterprise Web • Marketing Tech • News + Entertainment • Software • Generative AI
In-Office
San Francisco, CA, USA
14 Employees
175K-275K Annually

Hedra Logo Hedra

Scientist

Consumer Web • Digital Media • Enterprise Web • Marketing Tech • News + Entertainment • Software • Generative AI
In-Office
San Francisco, CA, USA
14 Employees
200K-325K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account