LLM Inference Engineer

Posted 22 Days Ago
2 Locations
In-Office or Remote
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Automation
The Role
Design, build, and maintain high-traffic production LLM serving systems. Optimize throughput, latency, and cost for open-source language models by tuning inference engines and GPU execution stacks. Troubleshoot and improve inference performance, and collaborate to scale efficient, reliable inference infrastructure.
Summary Generated by Built In

Locations: San Francisco or Remote

About The Role

The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.

We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.

What You'll Be Doing

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.

What We're Looking For

  • Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
  • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.

We'd Love If You Have

  • Experience with Trusted Execution Environments (TEE).
  • Active contributor to open-source LLM inference engines.

Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.

Skills Required

  • Hands-on experience in LLM inference and debugging/optimizing inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of GPU architectures and experience exploiting them with PyTorch, Triton, CuTe, CUDA.
  • Proven track record designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.
  • Experience with Trusted Execution Environments (TEE).
  • Active contribution to open-source LLM inference engines.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
9 Employees
Year Founded: 2017

What We Do

NEAR AI is an artificial intelligence research, engineering, and product development company committed to building an AI future owned by everyone. Founded by AI pioneer and former Google Deepmind researcher Illia Polosukhin, NEAR AI’s verifiable private inference infrastructure empowers developers and enterprises to deploy AI models with full control over their data. With hardware-backed private inference via a simple API, NEAR AI Cloud runs sensitive AI workloads securely and at scale, from privacy-critical consumer interactions to autonomous systems and critical infrastructure. NEAR AI Private Chat brings the same guarantees to users’ everyday questions and research. Serving over 100 million users across platforms such as Brave Nightly and OpenMind, NEAR AI is proven infrastructure for transforming sensitive data into safe intelligence and advancing a user-owned AI future. Learn more at https://near.ai/.

Similar Jobs

UL Solutions Logo UL Solutions

Field Business Manager - Northwest Region

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
4 Locations
15000 Employees
128K-145K Annually

UL Solutions Logo UL Solutions

Marketing Associate

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
United States
15000 Employees
55K-73K Annually

Coupa Logo Coupa

Americas Lead, Partner Success Manager - 11653

Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Remote
US
3000 Employees
208K-291K Annually

Wipfli Logo Wipfli

Consultant

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
3000 Employees
88K-118K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account