MTS 2, AI Platform

Posted 24 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
eCommerce • Retail
The Role
Own production LLM inference systems from model handoff through serving, optimizing latency, throughput, GPU utilization, memory, and cost. Operate heterogeneous GPU fleets, tune vLLM and Triton-based runtimes, build benchmarking and observability tools, improve reliability, and respond to incidents. Apply inference optimizations such as quantization, batching, KV-cache management, and runtime improvements. Partner with data science and product teams to establish performance and availability goals.
Summary Generated by Built In

At eBay, we're more than a global ecommerce leader — we’re changing the way the world shops and sells. Our platform empowers millions of buyers and sellers in more than 190 markets around the world. We’re committed to pushing boundaries and leaving our mark as we reinvent the future of ecommerce for enthusiasts.

Our customers are our compass, authenticity thrives, bold ideas are welcome, and everyone can bring their unique selves to work — every day. We're in this together, sustaining the future of our customers, our company, and our planet.

Join a team of passionate thinkers, innovators, and dreamers — and help us connect people and build communities to create economic opportunity for all.

As an LLM Inference Engineer on our AI Platform team, you’ll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable—the “last mile” from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps, and requires strong intuition for how model architecture, runtimes, and hardware interact.
What You’ll Do

  • Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.

  • Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.

  • Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e.g., schedulers, KV cache, batching, memory).

  • Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.

  • Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.

  • Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e.g., quantization, paging, kernel/runtimes improvements).

  • Partner with cross-functional teams: Work with data science and product teams to translate business requirements into performance and availability SLOs.

What We’re Looking For

  • 5+ years of strong development experience

  • Experience deploying and operating LLM inference services in production.

  • Strong production coding skills in Python plus Go or Rust (systems-level implementation and debugging).

  • Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang  (and/or TensorRT).

  • Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.

  • Solid understanding of LLM inference and optimization techniques: continuous batching, KV cache management, quantization, speculative decoding (nice-to-have), etc.

  • 3+ years hands-on experience in performance optimization and systems programming for AI/ML workloads.

  • Demonstrated ability to deliver measurable production improvements (e.g., 2X throughput, lower p95/p99 latency, reduced GPU cost).

  • Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.

  • Demonstrated proficiency in applying autonomous AI coding agents to speed up software delivery pipelines. This includes advanced prompting and careful human-in-the-loop code review to improve development speed and code accuracy.

Additional Details

eBay is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, sexual orientation, gender identity, veteran status, and disability, or other legally protected status. If you have a need that requires accommodation, please contact us at [email protected]. We will make every effort to respond to your request for accommodation as soon as possible. View our accessibility statement to learn more about eBay's commitment to ensuring digital accessibility for people with disabilities.


We use cookies to enhance your experience and may use AI tools for administrative tasks in the hiring process. To learn how we handle your personal data and use AI responsibly, please visit our Talent Privacy Notice, Privacy Center, and AI Hiring Guidelines.

Skills Required

  • 5+ years of strong software development experience
  • Experience deploying and operating LLM inference services in production
  • Strong production coding skills in Python plus Go or Rust
  • Experience with PyTorch and LLM runtimes such as vLLM, SGLang, or TensorRT
  • Knowledge of GPU architecture and performance, including profiling and memory tradeoffs
  • Understanding of LLM inference optimization techniques, including continuous batching, KV-cache management, and quantization
  • 3+ years of hands-on performance optimization and systems programming for AI/ML workloads
  • Demonstrated ability to deliver measurable production improvements in throughput, latency, or GPU cost
  • Root-cause analysis skills across models, runtimes, networking, and infrastructure
  • Proficiency applying autonomous AI coding agents, advanced prompting, and human-in-the-loop code review
  • CUDA or kernel programming experience
  • Experience with speculative decoding

eBay Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about eBay and has not been reviewed or approved by eBay.

  • Healthcare Strength — Medical, dental, and vision coverage begin on the date of hire, complemented by mental health resources, an EAP, disability coverage, FSAs, and wellness initiatives.
  • Leave & Time Off Breadth — A robust mix of PTO, paid holidays, flexible work styles, parental leave, and a four‑week paid sabbatical after five years supports work‑life balance.
  • Equity Value & Accessibility — Total compensation commonly includes stock components, with access to an employee stock purchase plan at a discount and stock awards enhancing overall packages.

eBay Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, CA
26,035 Employees

What We Do

eBay Inc. is a global commerce leader that connects millions of buyers and sellers around the world. We exist to enable economic opportunity for individuals, entrepreneurs, businesses and organizations of all sizes. Our portfolio of brands includes eBay Marketplace and eBay Classifieds Group, operating in 190 markets around the world. We offer sellers the ability to grow a business with little barrier to entry regardless of size, background or geographic location. We never compete with our sellers. We win when our sellers succeed. Buyers who shop on our Marketplace and Classifieds platforms enjoy a highly personalized experience with an unparalleled selection at great value.

Similar Jobs

Atlassian Logo Atlassian

Senior Engineering Manager

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
Bengaluru, Bengaluru Urban, Karnataka, IND
11000 Employees

Atlassian Logo Atlassian

Senior Engineering Manager

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
Bengaluru, Bengaluru Urban, Karnataka, IND
11000 Employees

Atlassian Logo Atlassian

Senior Engineering Manager

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
In-Office or Remote
Bengaluru, Bengaluru Urban, Karnataka, IND
11000 Employees

Coursera + Udemy  Logo Coursera + Udemy

Accounts Payable Specialist

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
India
1500 Employees

Similar Companies Hiring

Tastewise Thumbnail
Artificial Intelligence • Big Data • Food • Retail • Software • Generative AI • Big Data Analytics
NYC, NYC
120 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account