Software Engineer, Inference Runtime

Posted 3 Days Ago
Be an Early Applicant
New York City, NY, USA
Hybrid
150K-350K Annually
Senior level
Artificial Intelligence • Information Technology • Software
The Role
Develop and optimize inference runtimes for on-device and cloud deployments. Integrate inference engines, bring up new models and modalities, improve latency, throughput, memory, and reliability across CPU and GPU targets, build batching/scheduling/caching/distributed execution capabilities, benchmark and diagnose performance and correctness, and contribute upstream to open-source projects.
Summary Generated by Built In

LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family.

As a team, we work with high technical intensity and personal responsibility. We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software.

The Role

We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on.

Qualifications

  • Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure

  • Strong programming ability in Python and C++

  • Deep understanding of transformer architectures and the mechanics of model inference

  • Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement

  • Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

  • Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution

  • Takes personal responsibility for the correctness and performance of their work

Bonus Qualifications

  • Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

Responsibilities

  • Maintain and push forward our inference stack on-device and in the cloud

  • Bring up new model architectures and multimodal models

  • Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes

  • Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution

  • Benchmark and diagnose correctness and performance problems across the inference stack

  • Contribute upstream to open-source projects such as llama.cpp and MLX

Benefits

  • Competitive salary and equity grants

  • Great medical, vision, dental healthcare plans

  • Catered team lunch / expensed dinners in the office

  • Flexible PTO

  • Flexible WFH

  • Sun-drenched office in SoHo in NYC

Skills Required

  • Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure
  • Strong programming ability in Python and C++
  • Deep understanding of transformer architectures and the mechanics of model inference
  • Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement
  • Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM
  • Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution
  • Past contributions to open-source inference runtime projects (e.g., llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Brooklyn, NY
44 Employees
Year Founded: 2023

What We Do

Download and run local LLMs on your computer 👾 https://lmstudio.ai/download

Similar Jobs

Anthropic Logo Anthropic

Software Engineer

Artificial Intelligence • Natural Language Processing • Generative AI
In-Office or Remote
3 Locations
2500 Employees
405K-485K Annually
Hybrid
4 Locations
350 Employees
165K-330K Annually

HiBob Logo HiBob

Business Development Representative

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
United States
1350 Employees
64K-64K Annually

NBCUniversal Logo NBCUniversal

Global Response & Intelligence Officer

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Hybrid
New York, NY, USA
58K-69K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account