Inference Systems Performance Architect

Posted 2 Days Ago
Be an Early Applicant
San Jose, CA, USA
In-Office
245K-325K Annually
Expert/Leader
Artificial Intelligence • Hardware • Machine Learning • Natural Language Processing • Software • Generative AI
SambaNova is the #1 platform for business AI.
The Role
Lead end-to-end performance for large-scale LLM inference: build reproducible workload capture and agentic benchmarking, develop performance modeling and simulation, create distributed profiling tooling, guide cross-functional trade-offs, mentor senior engineers, and represent SambaNova to customers and partners to meet SLOs and inform system/hardware planning.
Summary Generated by Built In

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.

SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.

About the role

As an Architect on the Inference Systems Performance team, you'll own the discipline of end-to-end performance for large-scale LLM inference at SambaNova, from how a request moves through tokenization, prefill, decode, and the fabric between them, to how an entire deployment is sized against customer SLOs. Inference-systems performance is a nascent field; the results of design choices are being discovered daily rather than inherited from a mature craft, and this role exists to bring rigor to that frontier.

The work spans two coupled pillars.

The first is reproducible workload capture and benchmarking -- building faithful, replayable representations of real and increasingly agentic traffic, so that what we measure reflects production rather than an artifact of a naive load script. 

The second is performance modeling and simulation - analytic and simulation models that turn measurement into a "what-if" capability, letting us reason about configurations and hardware that do not exist yet.

Together these feed both today's serving optimization and the next generation of system planning.

The technical frontier you'll help define is heterogeneous, disaggregated inference - GPU on prefill, the RDU on decode - which explores hard problems across networking, storage, prompt caching, and tail-latency-bound data movement.

You will be the go-to person for inference-systems performance across SambaNova, and a resource the entire organization relies on to answer "how fast can this go, and what will it take."

Responsibilities

  • Define and drive the technical strategy for inference-systems performance including workload capture, benchmarking, modeling, and simulation, while developing and architecture that enables many potential futures
  • Build the workload-capture and agentic-benchmarking capability - capture representative production traffic and enforce the discipline of interrogating results, spotting artificial contention or misleadingly high cache-hit rates that never occur in real use
  • Own the performance-modeling and simulation practice - models that predict how a configuration change moves the output, informing capacity planning against customer SLOs and next-generation system and hardware planning
  • Attack the end-to-end profiling gap - drive tooling that produces accurate, actionable profiles of a distributed inference pipeline so bottlenecks can be localized across host, accelerator, and fabric
  • Serve as the senior technical voice across model-optimization, systems, hardware, and product, tying together multiple engineering activities and teams, and weighing trade-offs of reliability, scalability, operational cost, and ease of adoption
  • Act as a resource for the entire organization including representing SambaNova's performance story to customers and partners
  • Mentor and multiply by raising the capability of principal and senior engineers, building the systems, tools, and patterns that make everyone more productive
  • Drive the resolution of the most ambiguous, novel challenges that span organizational boundaries or have no established answer in the field yet

Required qualifications

  • 12+ years of experience in performance engineering, with a demonstrated record of technical leadership on large-scale, complex systems
  • Deep expertise in end-to-end performance analysis of distributed systems with many moving parts and the ability to localize bottlenecks that others cannot
  • Proven command of realistic workload generation and simulation and of performance modeling, including calibrating models against real, variable workloads
  • Demonstrated ability to enter an unfamiliar domain and apply core performance methods with transferable discipline expertise 
  • Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction
  • Experience representing an organizations credibly to customers and partners
  • Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant impact on products or roadmap

Preferred qualifications

  • Direct experience with LLM inference serving - continuous batching, prompt/KV caching, prefill/decode disaggregation, tail-latency SLOs
  • Familiarity with inference simulation frameworks or agentic benchmarking efforts
  • A public technical voice - talks, writing, or community presence on systems performance

Base Salary Range:

Base Pay Range
$245,000$325,000 USD

Submission Guidelines
Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified. 

EEO Policy
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.

Benefits Summary for US-Based, Full-Time Employment Positions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.

Skills Required

  • 12+ years of experience in performance engineering with technical leadership on large-scale, complex systems
  • Deep expertise in end-to-end performance analysis of distributed systems and ability to localize bottlenecks
  • Proven command of realistic workload generation, simulation, and performance modeling, including calibrating models against variable workloads
  • Ability to enter unfamiliar domains and apply core performance methods with disciplined, transferable expertise
  • Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction
  • Experience representing an organization credibly to customers and partners
  • Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant product or roadmap impact
  • Direct experience with LLM inference serving (continuous batching, prompt/KV caching, prefill/decode disaggregation, tail-latency SLOs)
  • Familiarity with inference simulation frameworks or agentic benchmarking efforts
  • A public technical voice (talks, writing, or community presence on systems performance)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
500 Employees
Year Founded: 2017

What We Do

AI is changing the world and at SambaNova, we believe that you don’t need unlimited resources to take advantage of the most advanced, valuable AI capabilities - capabilities that are helping organizations explore the universe, find cures for cancer, and giving companies access to insights that provide a competitive edge. We deliver the world’s fastest and only complete AI solution for enterprises and governments with world-record inference performance and accuracy. Powered by the SambaNova SN40L Reconfigurable Dataflow Unit (RDU), organizations can build a technology backbone for the next decade of AI innovation with SambaNova Suite. Our fully integrated hardware-software system, DataScale®, enables organizations to train, fine-tune, and deploy the most demanding AI workloads using the largest and most challenging models. Most recently, with the launch of our newest offering, SambaNova Cloud, developers can supercharge AI-powered applications on Llama 3.2 models. SambaNova was founded in 2017 in Palo Alto, California, by a group of industry luminaries, business leaders, and world-class innovators who understand AI. Today, we’ve built an incredibly smart and motivated team dedicated to making a lasting impact on the industry and equipping our customers to thrive in the new era of AI.

Why Work With Us

As a talent first company, we aim to hire the greatest and most innovative minds in the industry- driving the next generation of AI computing where no barrier is too high and the possibilities are truly limitless. We encourage our peers to take risks and take the initiative to make a lasting impact on the AI and ML industries.

Gallery

Gallery

Similar Jobs

Enverus Logo Enverus

Owner Relations Agent - 25270

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
3 Locations
1800 Employees
43K-58K Annually

Enverus Logo Enverus

Consultant

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
5 Locations
1800 Employees
120K-135K Annually
Easy Apply
In-Office
San Francisco, CA, USA
70 Employees
230K-330K Annually

Mochi Health Logo Mochi Health

Designer

Healthtech • Telehealth
Easy Apply
In-Office
San Francisco, CA, USA
70 Employees
150K-225K Annually

Similar Companies Hiring

Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account