Principal Engineer, Inference Memory and Storage Systems

Posted Yesterday
Be an Early Applicant
Seattle, WA, USA
Hybrid
250K-312K Annually
Expert/Leader
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
DigitalOcean is the Inference Cloud built for production AI.
The Role
Own the technical vision and roadmap for DigitalOcean’s unified memory and storage layer supporting large-scale LLM inference. Design multi-tier systems across GPU memory, host memory, NVMe, and remote storage; integrate with vLLM, SGLang, and TensorRT-LLM; define KV-cache transfer, eviction, admission, and routing protocols; optimize performance and unit economics; and lead cross-functional architecture, mentorship, open-source contributions, and customer-facing technical initiatives.
Summary Generated by Built In

Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here.  We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world. 

We are looking for a Principal Engineer to own the technical vision and roadmap for the memory and storage layer underneath DigitalOcean's inference platform.

DigitalOcean is the Inference Cloud. The Inference Platform team builds the serving stack that runs frontier open models on our GPU fleet at production scale—model orchestration, disaggregated prefill and decode, request routing, and the engine integrations (vLLM, SGLang, TensorRT-LLM) that make it all go. As context lengths grow and agentic traffic patterns become the norm, the single biggest lever on cost, TTFT, and throughput is where the KV cache lives and how efficiently we can move it.

That layer is currently a collection of good decisions made independently. This role exists to make it one system: a unified, multi-tier memory and storage substrate spanning GPU HBM, host DRAM, local NVMe, and remote object storage, exposed through a coherent interface to every serving engine and router we run. The selected candidate will set direction across teams, contribute meaningfully to the open source projects we depend on, and translate hard systems work into unit economics our customers feel.

What You'll Be Doing:
  • Defining and evolving a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, local NVMe tiers, and remote object storage for large-scale LLM inference
  • Architecting deep integrations with the serving engines we run in production (vLLM, SGLang, TensorRT-LLM), focused on KV cache offload, reuse, eviction policy, and cross-node sharing
  • Designing the interfaces and protocols behind disaggregated prefill/decode, peer-to-peer KV cache transfer, and cache-aware routing—including how the router, the engine, and the cache tier agree on what is resident and where
  • Owning the eviction and admission story across tiers (LRU and its successors, cost-aware policies, pluggable backends) and the metrics that prove those policies work under real traffic
  • Partnering with our GPU infrastructure, networking, and platform teams to exploit GPUDirect, RDMA, NVMe-oF, and NVLink for low-latency cache access across heterogeneous accelerator pools
  • Driving the unit economics: modeling and validating the effect of cache hit ratio, prefix caching, and tiering on $/token and $/GPU-hr, and making the tradeoffs legible to product and pricing partners
  • Setting technical direction and raising the bar through design review, mentorship of senior and staff engineers, and sponsorship of the initiatives that follow from this roadmap
  • Representing DigitalOcean externally—upstream contributions, conference talks, and customer-facing technical deep dives
What You'll Add to DigitalOcean:
  • 15+ years building large-scale distributed systems, high-performance storage, or ML systems infrastructure, with a track record of delivering and operating production services
  • Deep fluency in memory hierarchies (GPU HBM, host DRAM, NVMe, remote/object storage) and experience designing systems that span tiers for both performance and cost
  • Experience with distributed caching or key-value systems, particularly designs optimized for low latency under high concurrency
  • Hands-on work with networked I/O and RDMA / NVMe-oF / NVLink-class technologies, and familiarity with aggregated and disaggregated deployment topologies for AI clusters
  • Strong systems programming skills in C/C++, Go, Rust, or Python, and comfort reading and modifying serving-engine internals
  • Rigor in profiling and optimization across CPU, GPU, memory, and network, using measurement to drive architectural decisions and to validate improvements in TTFT, ITL, and throughput
  • Excellent written and verbal communication, and a history of leading cross-functional efforts with product, infrastructure, and customer-facing teams
Bonus:
  • Contributions to open source LLM serving or inference infrastructure projects—vLLM, SGLang, llm-d, NVIDIA Dynamo, LMCache, or similar—especially on KV cache offload, compression, or reuse
  • Experience designing a unified memory or storage layer that presents a single logical KV or object model across GPU, host, SSD, and cloud tiers in a hyperscale or public cloud environment
  • Familiarity with Kubernetes-based GPU orchestration, including DRA, MIG/MPS partitioning, and gateway/inference-extension routing patterns
  • Publications or patents in LLM systems, memory-disaggregated architectures, RDMA-based data planes, or CDN-like caching systems for ML workloads
Compensation Range: 
  • $249,600 - $312,000

*This is a hybrid role

JR: 2026-8160

#LI-Hybrid

Why You’ll Like Working for DigitalOcean
  • We innovate with purpose. You’ll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of responsibility for customers, products, employees, and decisions.
  • We prioritize career development. At DO, you’ll do the best work of your career. You will work with some of the smartest and most interesting people in the industry. We are a high-performance organization that will always challenge you to think big. Our organizational development team will provide you with resources to ensure you keep growing. We provide employees with reimbursement for relevant conferences, training, and education. All employees have access to LinkedIn Learning's 10,000+ courses to support their continued growth and development.
  • We care about your well-being. Regardless of your location, we will provide you with a competitive array of benefits to support you from our Employee Assistance Program to Local Employee Meetups to flexible time off policy, to name a few. While the philosophy around our benefits is the same worldwide, specific benefits may vary based on local regulations and preferences.
  • We reward our employees. The salary range for this position is based on market data, relevant years of experience, and skills. You may qualify for a bonus in addition to base salary; bonus amounts are determined based on company and individual performance. We also provide equity compensation to eligible employees, including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
  • DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.

Application Limit: You may apply to a maximum of 3 positions within any 180-day period. This policy promotes better role-candidate matching and encourages thoughtful applications where your qualifications align most strongly.

Skills Required

  • 15+ years building large-scale distributed systems, high-performance storage, or ML systems infrastructure
  • Track record of delivering and operating production services
  • Deep fluency in GPU HBM, host DRAM, NVMe, and remote/object storage memory hierarchies
  • Experience designing multi-tier systems for performance and cost
  • Experience with distributed caching or key-value systems optimized for low latency and high concurrency
  • Hands-on experience with networked I/O and RDMA, NVMe-oF, or NVLink-class technologies
  • Familiarity with aggregated and disaggregated AI cluster deployment topologies
  • Strong systems programming skills in C/C++, Go, Rust, or Python
  • Ability to read and modify serving-engine internals
  • Experience profiling and optimizing CPU, GPU, memory, and network performance
  • Excellent written and verbal communication skills
  • History of leading cross-functional efforts with product, infrastructure, and customer-facing teams
  • Contributions to open-source LLM serving or inference infrastructure projects
  • Experience designing unified memory or storage layers across GPU, host, SSD, and cloud tiers
  • Familiarity with Kubernetes-based GPU orchestration, DRA, MIG/MPS, and gateway or inference-extension routing
  • Publications or patents in LLM systems, memory-disaggregated architectures, RDMA data planes, or ML caching systems

What the Team is Saying

DigitalOcean Compensation & Benefits Highlights

  • Healthcare Strength — Health coverage is described as comprehensive, spanning medical, dental, vision, and mental-health insurance. Feedback suggests this area stands out as a core strength of the package.
  • Equity Value & Accessibility — Equity participation is emphasized through performance grants and an Employee Stock Purchase Plan. Feedback suggests these programs add meaningful value alongside base pay.
  • Leave & Time Off Breadth — Time off is positioned as flexible or unlimited, with above-average parental leave highlighted across sources. Feedback suggests this flexibility supports work-life balance.

DigitalOcean Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Broomfield, CO
1,400 Employees
Year Founded: 2012

What We Do

DigitalOcean is the Inference Cloud — a full-stack, production-ready cloud platform built to run AI applications with predictable performance, sustainable economics, and radically simpler operations at scale. We are built for teams turning AI into real products — not just training models. Our advantage is not fewer features, but fewer failure modes when operating AI at scale — combining minimal operational overhead, predictable cost efficiency, and a full-stack cloud that works as a system. Hyperscalers are broad by design. Neoclouds are infrastructure-first. DigitalOcean is inference-first — with a real cloud underneath. It combines inference-optimized compute, managed inference software, and integrated cloud capabilities that reduce operational burden for teams running real workloads. Inference is the foundation—not the boundary. Everything else builds on top of it.

Why Work With Us

At DO, we do career-defining work. We innovate with AI and build cutting-edge tech. Our rewards to match that intensity - to motivate you, recognize your impact, and give you what you need to thrive. If you have a growth mindset, like to think big and bold, and are energized by the fast-paced environment, you'll find your place here.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery

DigitalOcean Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

We commit to both remote work and in-person collaboration. These ways of working are dependent on specific roles and are mutually agreed upon by employees. In the US, we are mainly remote. In our APAC locations, we have a hybrid in-office approach.

Typical time on-site: Not Specified
Company Office Image
HQBroomfield, CO
Company Office Image
Seattle, WA
Company Office Image
Hyderabad, Telangana
Learn more

Similar Jobs

DigitalOcean Logo DigitalOcean

Principal Engineer

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Hybrid
Seattle, WA, USA
1400 Employees
250K-312K Annually

DigitalOcean Logo DigitalOcean

Principal Engineer

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Hybrid
Seattle, WA, USA
1400 Employees
250K-312K Annually

DigitalOcean Logo DigitalOcean

Senior People Business Partner

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Hybrid
Seattle, WA, USA
1400 Employees
95K-119K Annually

DigitalOcean Logo DigitalOcean

Scientist

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Hybrid
Seattle, WA, USA
1400 Employees
139K-174K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account