LLM Inference Engineer

Reposted 2 Hours Ago
2 Locations
In-Office or Remote
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Automation
The Role
Design, build, and maintain high-traffic production LLM serving systems. Optimize throughput, latency, and cost for open-source language models by tuning inference engines and GPU execution stacks. Troubleshoot and improve inference performance, and collaborate to scale efficient, reliable inference infrastructure.
Summary Generated by Built In

Locations: San Francisco or Remote

About The Role

The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.

We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.

What You'll Be Doing

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.

What We're Looking For

  • Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
  • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.

We'd Love If You Have

  • Experience with Trusted Execution Environments (TEE).
  • Active contributor to open-source LLM inference engines.

Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.

Skills Required

  • Hands-on experience in LLM inference and debugging/optimizing inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of GPU architectures and experience exploiting them with PyTorch, Triton, CuTe, CUDA.
  • Proven track record designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.
  • Experience with Trusted Execution Environments (TEE).
  • Active contribution to open-source LLM inference engines.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
9 Employees
Year Founded: 2017

What We Do

NEAR AI is an artificial intelligence research, engineering, and product development company committed to building an AI future owned by everyone. Founded by AI pioneer and former Google Deepmind researcher Illia Polosukhin, NEAR AI’s verifiable private inference infrastructure empowers developers and enterprises to deploy AI models with full control over their data. With hardware-backed private inference via a simple API, NEAR AI Cloud runs sensitive AI workloads securely and at scale, from privacy-critical consumer interactions to autonomous systems and critical infrastructure. NEAR AI Private Chat brings the same guarantees to users’ everyday questions and research. Serving over 100 million users across platforms such as Brave Nightly and OpenMind, NEAR AI is proven infrastructure for transforming sensitive data into safe intelligence and advancing a user-owned AI future. Learn more at https://near.ai/.

Similar Jobs

Motive Logo Motive

Lead, Safety and Compliance Strategy (Remote USA)

Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation
Easy Apply
Remote
United States
4000 Employees
135K-170K Annually

Samsara Logo Samsara

Sales Operations Admin I

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
79K-134K Annually

UL Solutions Logo UL Solutions

Sales Executive

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
United States
15000 Employees
120K-202K Annually

MetLife Logo MetLife

Care Coordinator Long Term Care - 9/28/26 - 19472

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
50K-58K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account