Software Engineer, Runtime

Posted Yesterday
Be an Early Applicant
Palo Alto, CA, USA
In-Office
Entry level
Artificial Intelligence • Machine Learning • Software • Generative AI
The Role
Develop and optimize Ollama’s local runtime for open-model inference across macOS, Linux, and Windows. Responsibilities include model loading, scheduling, memory management, quantization, GPU backends, model architecture integrations, and performance improvements. The role involves Go and C/C++ development, close-to-the-metal systems work, profiling, hardware optimization, open-source collaboration, and partnerships with model labs and hardware vendors.
Summary Generated by Built In

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll work on the heart of Ollama — the local runtime that runs open models on developers' own machines. It loads models, manages memory, drives GPU acceleration across NVIDIA, AMD, Intel, Qualcomm, and Apple Silicon (including our MLX integration), and makes all of it feel instant. You'll work in Go and C/C++ and touch the model formats and inference engines underneath, shipping to macOS, Linux, and Windows across an enormous range of hardware.

What you'll do
  • Make open models run fast and reliably on consumer and enterprise hardware — from a MacBook Pro to server-grade NVIDIA GPUs.

  • Own pieces of the runtime: model loading & scheduling memory management, quantization, GPU hardware backends.

  • Integrate new model architectures and quantization formats so the latest open models work on day one.

  • Improve cold-start, time-to-first-token, and throughput

  • Partner with model labs and hardware vendors on early access and deep integrations.

  • Ship in the open: Ollama is open source, and you'll work with the community

Example projects
  • Add support for a new model family end-to-end — format parsing, weights loading, and the defaults that make it useful out of the box.

  • Cut cold-start for a popular model in half by streaming weights and lazy-loading layers.

  • Land a new quantization format so a 70B model runs on a single consumer GPU.

  • Wire up a new GPU backend and find a 2x throughput win with kernel selection and memory tuning.

  • Improve the "Auto" experience — picking the right model and settings for a machine's hardware without the user thinking about it.

You may be a fit if
  • You have strong systems fundamentals and are comfortable in Go, C, or C++

  • You've worked close to the metal — GPU compute, inference, game engines, databases, OS, or networking.

  • You care about performance and have profiled and optimized real workloads.

  • You're comfortable shipping to millions of users and handling the long tail of hardware and OS combinations.

  • Bonus: experience with model quantization, GPU programming (CUDA/Metal/SYCL), Apple MLX

Skills Required

  • Strong systems fundamentals
  • Proficiency in Go, C, or C++
  • Experience with close-to-the-metal software such as GPU compute, inference, game engines, databases, operating systems, or networking
  • Experience profiling and optimizing real workloads
  • Ability to ship software to millions of users across varied hardware and operating systems
  • Experience with model quantization
  • Experience with GPU programming, including CUDA, Metal, or SYCL
  • Experience with Apple MLX
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
64 Employees
Year Founded: 2023

What We Do

Ollama develops tools and infrastructure for running open-weight artificial intelligence models locally and in the cloud. Its software helps developers launch and interact with models through a command-line interface, APIs, integrations, and cloud services, while supporting offline use and emphasizing data privacy. The company’s mission is to make open models accessible and practical for developers’ workflows across regions and environments.

Similar Jobs

Hybrid
Santa Clara, CA, USA
471 Employees
130K-220K Annually

OpenAI Logo OpenAI

Software Engineer

Artificial Intelligence • Machine Learning • Generative AI
Hybrid
San Francisco, CA, USA
4500 Employees
266K-445K Annually

Loft Orbital Logo Loft Orbital

Software Engineer

Aerospace • Defense
In-Office
2 Locations
300 Employees
144K-198K Annually

crewAI Logo crewAI

Software Engineer

Artificial Intelligence • Software
In-Office
San Francisco, CA, USA
7 Employees

Similar Companies Hiring

Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account