Senior Software Engineering Manager – KV Cache Platform

Posted 4 Days Ago
Be an Early Applicant
Santa Clara, Tarapacá, Amazonas, COL
Hybrid
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
The Role
Leads geographically distributed software engineering teams building the KV Cache Platform for large-scale LLM inference. Owns technical strategy, architecture, roadmap execution, customer deployments, production readiness, and engineering quality. Partners with product, sales, customer engineering, NVIDIA, and strategic partners while driving distributed systems development involving GPU optimization, caching, storage, networking, and AI infrastructure. Builds engineering talent and oversees releases, escalations, testing, observability, and operational excellence.
Summary Generated by Built In

DDN is seeking a Senior Software Engineering Manager to lead the engineering organization responsible for our KV Cache Platform—a distributed memory and storage platform that accelerates large-scale LLM inference across GPU clusters.

In this role, you will lead geographically distributed engineering teams responsible for building highly scalable, low-latency distributed systems that power AI inference. You will define the technical vision and execution strategy for the platform while partnering closely with Product Management, Sales, Customer Engineering, NVIDIA, and executive leadership to deliver innovative AI infrastructure that meets customer needs and supports DDN's long-term product strategy.

 

This is a highly visible leadership role with responsibility for engineering execution, customer success, roadmap delivery, and building a world-class engineering organization.

 
Responsibilities
  • Lead, mentor, and grow a geographically distributed team of software engineers and technical leaders, fostering a culture of technical excellence, innovation, ownership, and collaboration.

  • Define and execute the technical strategy and roadmap for the KV Cache Platform, ensuring scalability, reliability, security, and operational excellence.

  • Drive the architecture, development, and delivery of distributed systems supporting AI inference, GPU memory optimization, distributed caching, RDMA networking, GPUDirect Storage, NVIDIA BlueField DPUs, and emerging AI infrastructure technologies.

  • Partner closely with Product Management, Sales, Customer Engineering, NVIDIA, and strategic technology partners to prioritize customer requirements, drive proof-of-concepts (POCs), influence product direction, and successfully deliver customer deployments.

  • Own day-to-day engineering execution, including feature development, release planning, bug triage, production issues, customer escalations, and cross-functional execution to ensure timely, high-quality software delivery.

  • Establish engineering best practices for software quality, observability, automation, performance, testing, and production readiness.

  • Collaborate across engineering, infrastructure, and hardware teams to deliver scalable, production-ready AI infrastructure while developing future engineering leaders and driving continuous improvement.

Qualifications

Required
  • 15+ years of experience building distributed systems, cloud infrastructure, storage platforms, or AI infrastructure software.

  • 7+ years leading high-performing software engineering organizations, including geographically distributed teams.

  • Proven experience delivering large-scale distributed infrastructure products from architecture through production deployment.

  • Strong background in distributed systems, Linux, networking, performance engineering, and cloud-native architectures.

  • Hands-on programming experience with Go and Python; experience with C/C++ is a plus.

  • Demonstrated ability to lead cross-functional initiatives and influence technical direction across multiple organizations.

  • Experience building AI infrastructure, LLM serving platforms, distributed caching systems, or high-performance storage solutions.

  • Experience with technologies such as NVIDIA Dynamo, TensorRT-LLM, Triton, RDMA, GPUDirect Storage, BlueField DPUs, Kubernetes, or related AI infrastructure.

  • Background in HPC, distributed storage, networking, or enterprise infrastructure software.

  • Experience working directly with strategic customers, technology partners, OEMs, or hyperscalers to deliver enterprise AI solutions.

Skills Required

  • 15+ years of experience building distributed systems, cloud infrastructure, storage platforms, or AI infrastructure software
  • 7+ years leading high-performing software engineering organizations, including geographically distributed teams
  • Experience delivering large-scale distributed infrastructure products from architecture through production deployment
  • Strong background in distributed systems, Linux, networking, performance engineering, and cloud-native architectures
  • Hands-on programming experience with Go and Python
  • Experience with C or C++
  • Ability to lead cross-functional initiatives and influence technical direction across multiple organizations
  • Experience building AI infrastructure, LLM serving platforms, distributed caching systems, or high-performance storage solutions
  • Experience with NVIDIA Dynamo, TensorRT-LLM, Triton, RDMA, GPUDirect Storage, BlueField DPUs, Kubernetes, or related AI infrastructure
  • Background in HPC, distributed storage, networking, or enterprise infrastructure software
  • Experience working directly with strategic customers, technology partners, OEMs, or hyperscalers to deliver enterprise AI solutions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chatsworth, CA
706 Employees
Year Founded: 1998

What We Do

DDN is the world’s largest private data storage company and the leading provider of intelligent technology and infrastructure solutions for Enterprise At Scale, AI and analytics, HPC, government and academia customers. Through its DDN and Tintri divisions, the company delivers AI, Data Management software and hardware solutions, and unified analytics frameworks to solve complex business challenges for data-intensive, global organizations. DDN provides its enterprise customers with the most flexible, efficient and reliable data storage solutions for on-premises and multi-cloud environments at any scale. Over the last two decades, DDN has established itself as the data management provider of choice for over 11,000 enterprises, government, and public-sector customers, including many of the world’s leading financial services firms, life science organizations, manufacturing and energy companies, research facilities, and web and cloud service providers.

Similar Jobs

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Sales Support Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
14 Locations
16000 Employees
15-20 Hourly

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Temporary Sales Associate

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
14 Locations
16000 Employees
15-20 Hourly

Cloudflare Logo Cloudflare

Senior Customer Engineer, Bogotá, Colombia.

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
Colombia
4400 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sr. Sales Associate III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Remote or Hybrid
14 Locations
16000 Employees
15-20 Hourly

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account