Senior Inference Optimization Engineer - Dragonfly Portfolio

Posted 2 Days Ago
Be an Early Applicant
34 Locations
Remote or Hybrid
Senior level
Fintech • Software • Financial Services • Cryptocurrency
The Role
Optimize large-model inference at scale: improve throughput, reduce latency and cost per token, build benchmarking harnesses, tune parallelism and quantization strategies, implement load-balancing in routing, and evaluate custom kernels and emerging inference hardware.
Summary Generated by Built In
Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.

This is an application to join our talent network. This is not a listing for an internal role at Dragonfly.

We're actively sourcing for a Senior Inference Optimization Engineer for one of our portfolio companies building privacy-first consumer AI infrastructure. You'll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.

Location: Remote, USA (open to excellent candidates outside the USA)

What We’re Looking For:
  • 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
  • Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
  • Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
  • Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
  • GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
  • Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
  • Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions

About the role:
  • Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
  • Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Optimize multivariate inference load-balancing algorithms within the inference routing system
  • Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack

Even if you don't match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.

Process: 
  • We'll review your application and assess fit for this role.
  • If there's a match, we'll facilitate a warm introduction to the team.
  • If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.

Submit your information below, and we’ll reach out if there’s a potential fit.

Skills Required

  • 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
  • Hands-on experience with a production LLM inference engine (vLLM, SGLang) running at high volume
  • Experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA Graphs, torch.compile
  • Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
  • GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
  • Proficiency in Python, Rust, or Go
  • Proficiency in C++ and CUDA
  • Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Nanterre
152 Employees

What We Do

A cross-border crypto venture fund. Global from day one.

Similar Jobs

Deepgram Logo Deepgram

Account Executive

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
In-Office or Remote
28 Locations
150 Employees

Zapier Logo Zapier

Engineer, Applied AI

Artificial Intelligence • Productivity • Software • Automation
Remote
29 Locations
800 Employees
192K-287K Annually

Coinbase Logo Coinbase

Controller

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
26 Locations
4700 Employees
79K-88K Annually

Coinbase Logo Coinbase

Regional Threat Assessment Manager

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
26 Locations
4700 Employees
95K-106K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account