Principal Product Manager, Augmented Memory Grid (AMG)

Posted 9 Hours Ago
Be an Early Applicant
Hiring Remotely in U.S.
Remote
Expert/Leader
Artificial Intelligence • Big Data • Machine Learning
The Role
Own the AMG product roadmap for NeuralMesh: define KV-cache/prefix-cache offload, memory tiering, and integrations with inference engines and orchestration. Translate GPU/memory/network tradeoffs into product requirements, validate benchmarks with enterprise customers and GPU-cloud partners, partner with NVIDIA for certifications, and support sales with technical positioning to reduce cost-per-token and improve inference SLAs at scale.
Summary Generated by Built In

WEKA is architecting a new approach to the enterprise data stack built for the age of reasoning. NeuralMesh by WEKA sets the standard for agentic AI data infrastructure with a cloud- and AI-native software solution that can be deployed anywhere. It transforms legacy data silos into data pipelines that dramatically increase GPU utilization and make AI model training and inference, machine learning, and other compute-intensive workloads run faster, work more efficiently, and consume less energy.

WEKA is a pre-IPO, growth-stage company on a hyper-growth trajectory. We’ve raised $375M in capital with dozens of world-class venture capital and strategic investors. We help the world’s largest and most innovative enterprises and research organizations, including 12 of the Fortune 50, achieve discoveries, insights, and business outcomes faster and more sustainably. We’re passionate about solving our customers’ most complex data challenges to accelerate intelligent innovation and business value. If you share our passion, we invite you to join us on this exciting journey.

About the role

WEKA is looking for a Product Manager to own the roadmap and go-to-market for Augmented Memory Grid (AMG), part of the NeuralMesh platform. This is a deeply technical PM role sitting at the intersection of AI inference infrastructure, high-performance networking, and enterprise storage. You will work directly with engineering, GPU/inference partners (NVIDIA, hyperscalers, GPU clouds), and enterprise customers running large-scale LLM inference to define what AMG needs to do next.

Bring Your Expertise – and Your Passion

  • Leadership Skills: Strong leadership skills with a history of successfully leading cross-functional teams. Product Managers are expected to inspire and motivate team members to achieve ambitious goals while maintaining a collaborative and positive working environment. You understand how to influence without authority, and your recall of meaningful details supports verbal and written agility.
  • Strategic Vision: You are a strategic thinker who can develop and execute product strategies that align with market trends and customer needs, as well as think critically about existing strategies. You have a proven ability to translate strategic goals into actionable plans and deliver results.
  • Communication Skills: You have excellent communication and interpersonal skills, with the ability to articulate complex technical concepts to both technical and non-technical stakeholders. You are comfortable presenting product strategies and roadmaps to internal teams and external customers.
What you’ll do
  • Own the AMG product roadmap: KV-cache/prefix-cache offload, memory tiering, and integration with inference engines and orchestration layers (vLLM, NVIDIA Triton/TensorRT-LLM/NIM, Kubernetes-based serving).
  • Partner with engineering to define architecture trade-offs across GPU memory, networking (RDMA, GPUDirect, NVMe-oF), and distributed storage — translating inference performance bottlenecks (time-to-first-token, throughput, context length) into product requirements.
  • Work directly with enterprise customers and GPU cloud partners: Nebius, CoreWeave, TogetherAI, etc., running production inference workloads to gather requirements, validate benchmarks, and prioritize features that reduce cost-per-token and improve SLAs at scale.
  • Partner with NVIDIA and other silicon/inference-stack partners on joint roadmap and certification work.
  • Define and track benchmarks (TTFT, throughput, cache hit rate) that demonstrate AMG's value versus standard GPU-memory-only inference.
  • Support sales and field teams with technical positioning, competitive differentiation, and enterprise deal support.
Must-have qualifications
  • Inference ecosystem depth: hands-on product or engineering experience with LLM inference serving — vLLM, NVIDIA Triton/TensorRT-LLM/NIM, Ray Serve, or comparable — and fluency in concepts like KV-cache, prefix/context caching, quantization, and batching strategies.
  • Model & systems familiarity: working knowledge of how modern LLMs are served in production (context windows, multi-tenant serving, GPU scheduling) well enough to translate model-level constraints into infrastructure requirements.
  • Networking/infrastructure fluency: comfort with the fundamentals of high-performance networking and distributed systems — RDMA, GPUDirect Storage, NVMe-oF, or equivalent — and how they affect inference performance.
  • Enterprise customer experience: track record working directly with large enterprise accounts — requirements gathering, production deployments, SLAs — not solely self-serve/PLG products.
  • 10+ years of product management experience, ideally with some portion in infrastructure, ML platforms, or developer-facing technical products.
Nice-to-have
  • Prior experience at a GPU cloud, inference platform, or AI infrastructure startup.
  • Familiarity with storage systems (parallel/distributed file systems, object storage) in AI/ML pipelines.
  • Experience partnering directly with NVIDIA or other accelerator/silicon vendors.

The WEKA Way:
  • We are Accountable: We take full ownership, always–even when things don’t go as planned. We lead with integrity, show up with responsibility & ownership, and hold ourselves and each other to the highest standards.
  • We are Brave: We question the status quo, push boundaries, and take smart risks when needed. We welcome challenges and embrace debates as opportunities for growth, turning courage into fuel for innovation.
  • We are Collaborative: True collaboration isn’t only about working together. It’s about lifting one another up to succeed collectively. We are team-oriented and communicate with empathy and respect. We challenge each other and conduct positive conflict resolution. We are being transparent about our goals and results. And together, we’re unstoppable.
  • We are Customer Centric: Our customers are at the heart of everything we do. We actively listen and prioritize the success of our customers, and every decision we make is driven by how we can better serve, support, and empower them to succeed. When our customers win, we win.

Concerned that you don’t meet every qualification above?

Studies have shown that women and people of color may be less likely to apply for jobs if they don’t meet every qualification specified. At WEKA, we are committed to building a diverse, inclusive and authentic workplace. If you are excited about this position but are concerned that your past work experience doesn’t match up perfectly with the job description, we encourage you to apply anyway – you may be just the right candidate for this or other roles at WEKA.

 WEKA is an equal opportunity employer that prohibits discrimination and harassment of any kind. We provide equal opportunities to all employees and applicants for employment without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training.

Skills Required

  • Hands-on product or engineering experience with LLM inference serving (vLLM, NVIDIA Triton, TensorRT-LLM/NIM, Ray Serve, or comparable)
  • Working knowledge of modern LLM production serving (context windows, multi-tenant serving, GPU scheduling)
  • Fluency with high-performance networking and distributed systems (RDMA, GPUDirect Storage, NVMe-oF) and their impact on inference performance
  • Experience working directly with large enterprise accounts for requirements gathering, production deployments, and SLA management
  • 10+ years of product management experience, with a portion in infrastructure, ML platforms, or developer-facing technical products
  • Strong cross-functional leadership, strategic vision, and ability to communicate complex technical concepts to technical and non-technical stakeholders
  • Experience defining and tracking inference benchmarks (time-to-first-token, throughput, cache hit rate)
  • Prior experience at a GPU cloud, inference platform, or AI infrastructure startup
  • Familiarity with storage systems used in AI/ML pipelines (parallel/distributed file systems, object storage)
  • Experience partnering directly with NVIDIA or other accelerator/silicon vendors
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Campbell, CA
273 Employees
Year Founded: 2014

What We Do

Weka offers WekaFS, the modern file system that uniquely empowers organizations to solve the newest, biggest problems holding back innovation. Optimized for NVMe and the hybrid cloud, Weka handles the most demanding storage challenges in the most data-intensive technical computing environments, delivering truly epic performance at any scale. Its modern architecture unlocks the full capabilities of today’s data center, allowing businesses to maximize the value of their high-powered IT investments. Weka helps industry leaders reach breakthrough innovations and solve previously unsolvable problems. Try now at https://www.weka.io/

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
9 Locations
40000 Employees
45K-85K Annually

HiBob Logo HiBob

Customer Experience Specialist

HR Tech • Information Technology • Professional Services • Sales • Software
Remote or Hybrid
United States
1350 Employees
106K-135K Annually

Flywire Logo Flywire

Program Manager

Fintech • Payments • Software
Remote or Hybrid
MA, USA
1200 Employees
115K-150K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Counsel

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
2 Locations
40000 Employees
150K-209K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account