Software Engineer, Cloud

Posted 3 Days Ago
Palo Alto, CA, USA
In-Office
Entry level
Artificial Intelligence • Machine Learning • Software • Generative AI
The Role
Build and scale Ollama’s cloud inference platform, including high-throughput serving, GPU and regional workload routing, multi-tenant infrastructure, quotas, metering, billing, tiering, reliability, observability, and cost controls. The role requires production ownership of distributed systems and familiarity with Kubernetes, GPU scheduling, inference infrastructure, reliability practices, SLOs, and capacity planning.
Summary Generated by Built In

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.

What you'll do
  • Build and scale the inference platform that serves every request from ollama.com.

  • Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.

  • Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.

  • Build the reliability, observability, and cost controls for our team and customers

You may be a fit if
  • You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.

  • You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.

  • You've worked with Kubernetes, GPU scheduling, or inference infrastructure.

  • You think in terms of reliability, SLOs, and honest capacity planning.

  • Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

Skills Required

  • Deep experience with high-throughput, low-latency distributed systems, such as inference serving, traffic routing, real-time data pipelines, or large-scale APIs
  • Experience owning a production service end-to-end
  • Ability to evaluate cost and performance tradeoffs at scale
  • Experience with Kubernetes, GPU scheduling, or inference infrastructure
  • Understanding of reliability, SLOs, and capacity planning
  • Experience building an inference platform, GPU fleet management, or billing and metering for an AI service
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
64 Employees
Year Founded: 2023

What We Do

Ollama develops tools and infrastructure for running open-weight artificial intelligence models locally and in the cloud. Its software helps developers launch and interact with models through a command-line interface, APIs, integrations, and cloud services, while supporting offline use and emphasizing data privacy. The company’s mission is to make open models accessible and practical for developers’ workflows across regions and environments.

Similar Jobs

Hybrid
3 Locations
289097 Employees

Poshmark Logo Poshmark

Software Engineer

Consumer Web • eCommerce • Fashion • Retail
In-Office
Redwood City, CA, USA
850 Employees
108K-205K Annually
In-Office or Remote
2 Locations
501 Employees
118K-194K Annually

General Motors Logo General Motors

Senior Software Engineer

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
4 Locations
165000 Employees
129K-260K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account