Principal Software Engineer, Inference

Posted 19 Days Ago
Be an Early Applicant
3 Locations
In-Office
152K-349K Annually
Expert/Leader
Artificial Intelligence • Cloud • Information Technology • Consulting
The Role
Lead the architecture and technical direction of HPE’s LLM inference runtime and Kubernetes orchestration platform. Own engine integration, continuous batching, KV cache management, quantization, distributed inference, GPU scheduling, autoscaling, and performance optimization. Evaluate emerging inference technologies, mentor engineers, lead architecture reviews, and communicate technical strategy to executives. The role targets low latency, high throughput, and efficient GPU utilization across enterprise, air-gapped, and sovereign environments.
Summary Generated by Built In
Principal Software Engineer, Inference

  

This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description:

   

HPE's Private Cloud AI organization is seeking a Principal Software Engineer to lead the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The principal engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will define the architecture of that runtime – engine integration, batching, KV cache management, and distributed execution – together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered.

Responsibilities

·       Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution

·       Partner with inference performance engineering teams, with accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency

·       Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, KV cache offload across GPU memory, host memory, and RDMA-attached storage

·       Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether each runtime is adopted, developed in-house, or declined

·       Define the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling

·       Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences

Knowledge and Skills

Required

·       Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals

·       Comprehensive understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding

·       Tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them

·       Expert level proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling

·       Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight

·       Experience with debugging/profiling multi-tier application workloads such as RAG, Agents, etc

·       Excellent analytical, debugging, and problem-solving abilities

Preferred

·       Upstream contribution to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe

·       Disaggregated prefill/decode serving, or KV cache offload and reuse at scale

·       RDMA, GPUDirect Storage, InfiniBand, or RoCE

·       MIG, fractional GPU allocation, and multi-tenant GPU isolation

·       On-premises, air-gapped, or regulated enterprise software delivery

Experience and Education

·       Minimum of 12 years of experience in Software Engineering, including +1 years working directly on LLM inference runtimes or production model serving

·       Degree in Computer Science or related field

Accessibility


HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here.


Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.

What We Can Offer You:

Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected:

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

#unitedstates

Job:

Engineering

Job Level:

TCP_05

    

"The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.
– United States of America: Annual Salary USD 160,000 - 303,000 in Colorado // 152,000 - 349,000 in North Carolina & Texas
The listed salary range reflects base salary. Variable incentives may also be offered."

Information about employee benefits offered in the US can be found at https://myhperewards.com/main/new-hire-enrollment.html

The estimated job application period closure is December 30 2027; this timeline is provided for transparency and internal planning purposes.

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity.

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

   

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

   

Recruitment Fraud Alert

We have become aware of an increase in fraudulent recruitment activities in which individuals impersonate our company or authorized recruitment agencies to offer fake employment opportunities. These scams may occur through false websites, emails, social media, or chat-based applications and often aim to obtain personal information or money. Please note that Hewlett Packard Enterprise (HPE), its direct and indirect subsidiaries and affiliated companies, and its authorized recruitment agencies/vendors will never charge a candidate a registration fee, hiring fee, or any other fee in connection with its recruitment and hiring process. We also never request personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications.

All legitimate job opportunities will come through official company channels, and candidates are responsible for verifying the credentials of any third party claiming to represent the company. Any reliance on fraudulent communication is at the individual’s own risk, and HPE disclaims legal liability for any resulting damages. If you suspect recruitment fraud, do not share personal information or make any payments and report the incident to your local authorities immediately.

Skills Required

  • Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals
  • Comprehensive understanding of continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
  • Experience with tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and GPU interconnect characteristics
  • Expert-level proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling
  • Strong programming proficiency in Go and Python, with ability to read, debug, and profile C++ and CUDA using tools such as Nsight
  • Experience debugging and profiling multi-tier application workloads such as RAG and agent systems
  • Excellent analytical, debugging, and problem-solving abilities
  • Minimum 12 years of software engineering experience, including at least 1 year working directly on LLM inference runtimes or production model serving
  • Degree in Computer Science or a related field
  • Upstream contribution to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe
  • Experience with disaggregated prefill/decode serving or KV cache offload and reuse at scale
  • Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE
  • Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation
  • Experience delivering on-premises, air-gapped, or regulated enterprise software

Hewlett Packard Enterprise Compensation & Benefits Highlights

  • Parental & Family Support — Parental leave is described as industry-leading at six months fully paid, with additional transition options for new parents and support for adoption and fertility. Family-care resources such as backup childcare are also cited.
  • Leave & Time Off Breadth — Wellness Fridays, generous PTO, paid holidays, and paid volunteer time are highlighted, with flexible or unlimited time off reported in some groups. These programs are often associated with strong work-life balance.
  • Healthcare Strength — Medical, dental, vision, life and disability insurance, mental-health resources, and wellness programs are presented as comprehensive and widely available. Coverage breadth is frequently acknowledged as a core part of the package.

Hewlett Packard Enterprise Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Houston, TX
85,422 Employees
Year Founded: 2015

What We Do

In 1939, Bill Hewlett and Dave Packard, college friends turned business partners, started the original Silicon Valley startup in the space of a rented Palo Alto garage. Starting with audio oscillators, the friends built the foundation for a company that would grow to become a global leader in enterprise technology. More than 75 years later, our success is exemplified through our employees’ drive to advance ideas that bring meaningful innovations to life for our customers and partners around the globe. We are guided by our mission to help customers use technology to turn ideas into value, and empower them to transform industries, markets and lives. We simplify Hybrid IT, power the Intelligent Edge and provide the expertise to make it all happen.

Hewlett Packard Enterprise Offices

OnSite Workspace

Typical time on-site: None
HQHouston, TX
Heredia
SG
Alpharetta, US
Bangalore, Bangalore
Bengaluru, IN
Berkshire, GB
Durham, US
Fort Collins, US
Frisco, US
Galway, IE
Guadalajara, MX
Gurugram, Haryana
Kolkata, West Bengal
London, GB
Madrid, ES
New York, US
Puteaux, FR
Reading, GB
Riyadh, SA
Roseville, US
San Jose, US
Singapore, SG
Spring, US
Tokyo, JP
Learn more

Similar Jobs

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Senior Thermal Engineer (Houston, TX)

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
Spring, TX, USA
85422 Employees
109K-251K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Project Manager

Artificial Intelligence • Cloud • Information Technology • Consulting
Remote or Hybrid
5 Locations
85422 Employees
93K-214K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

HPC & AI Performance Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
2 Locations
85422 Employees
106K-243K Annually

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Hardware Mechanical Engineer II

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
3 Locations
85422 Employees
63K-145K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account