Software Engineer — Distributed LLM Inference Systems

Posted 11 Hours Ago
Be an Early Applicant
Shanghai, Shanghai Municipality, Shanghai, CHN
In-Office
Entry level
Artificial Intelligence • Cloud • Information Technology • Software
Creating world-changing technology that enriches the lives of every person on earth.
The Role
Design, implement, and optimize distributed LLM inference systems and framework components. Implement distributed algorithms, schedulers, KV cache management, and communication layers. Profile workloads to find bottlenecks, collaborate to improve latency, throughput, scalability, and contribute code, tests, and documentation to internal and open-source projects.
Summary Generated by Built In
Job Details:

Job Description: 

The Role and Impact:

  • As a Software Engineer on Intel’s Artificial Intelligence Frameworks team, you will contribute to designing, developing, and optimizing distributed inference systems for large language models.
  • Your day-to-day work will involve implementing distributed inference algorithms, optimizing model execution and communication, transforming neural network models, and developing software components that improve inference performance across diverse hardware architectures. You may work on areas such as disaggregated serving, request scheduling, KV cache management, parallel execution, and efficient communication between inference components.
  • By collaborating with researchers and engineers, you will play a key role in advancing Intel's AI capabilities and ensuring industry-leading solutions.

Business Group: Intel's Artificial Intelligence Frameworks team is dedicated to empowering transformative AI solutions by developing and optimizing software frameworks for machine learning and deep learning. This group works on enhancing the performance of AI applications across diverse computing hardware backends while contributing to open-source communities. As part of Intel, this team supports the mission to drive technological innovation and deliver impactful AI advancements globally.

Key Responsibilities

  • Design, develop, and optimize distributed LLM inference systems and related AI framework components.
  • Implement distributed algorithms, including model/data parallel frameworks and asynchronous communication for deep learning.
  • Develop and optimize components such as request schedulers, model workers, communication layers, and KV cache management mechanisms.
  • Profile distributed inference workloads to identify computation, communication, memory, and scheduling bottlenecks.
  • Collaborate with component teams to improve end-to-end latency, throughput, scalability, and resource utilization.
  • Contribute high-quality code, tests, and documentation to internal and open-source projects while following industry engineering standards.

Qualifications:

Minimum Qualifications –

  • Master’s degree in computer science, Artificial Intelligence, Software Engineering, or a related field, with 0-1 years of hands-on experience demonstrated through internships, academic projects, coursework, or training.
  • Proficiency in Python and modern C++ programming.
  • Foundational knowledge of deep learning and AI frameworks, such as PyTorch.
  • Experience debugging and optimizing software for performance.
  • Basic understanding of machine learning algorithms and techniques.
  • Strong problem-solving skills and the ability to learn unfamiliar systems quickly.

Preferred Qualifications

  • Experience or project exposure related to distributed LLM inference and serving.
  • Experience in contributing to open-source projects or collaborating within open-source ecosystems.
  • Understanding of LLM inference concepts such as prefill and decode, KV cache management, continuous batching, parallelism strategies, and disaggregated serving.
  • Familiarity with inference engines or serving frameworks such as vLLM, SGLang, TensorRT-LLM, or similar technologies.
  • Knowledge of large language models and inference optimization techniques.
  • Knowledge of AI Agent architecture and execution workflows, including tool calling, planning, memory, context management, and multi-agent coordination.
  • Effective communication skills, including fluency in written and spoken English. Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations.

Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations. Your contributions will shape the future of AI software and its real-world impact. Apply now and be a part of advancing transformative AI capabilities.

          

Job Type:College Grad

Shift:Shift 1 (China)

Primary Location: PRC, Shanghai

Additional Locations:

Posting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/A

Work Model for this Role

This role will require an on-site presence. * Job posting details (such as work model, location or time type) are subject to change.

*

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.

Skills Required

  • Master's degree in Computer Science, Artificial Intelligence, Software Engineering, or related field
  • 0-1 years hands-on experience (internships, projects, coursework)
  • Proficiency in Python
  • Proficiency in modern C++
  • Foundational knowledge of deep learning and AI frameworks such as PyTorch
  • Experience debugging and optimizing software for performance
  • Basic understanding of machine learning algorithms and techniques
  • Strong problem-solving skills and ability to learn unfamiliar systems quickly
  • On-site presence in Shanghai (PRC) / Shift 1 (China)
  • Experience or project exposure related to distributed LLM inference and serving
  • Experience contributing to or collaborating in open-source projects
  • Understanding of LLM inference concepts (prefill/decode, KV cache, continuous batching, parallelism, disaggregated serving)
  • Familiarity with inference engines/serving frameworks such as vLLM, SGLang, TensorRT-LLM
  • Knowledge of large language models and inference optimization techniques
  • Knowledge of AI Agent architecture and execution workflows (tool calling, planning, memory, multi-agent coordination)
  • Fluency in written and spoken English

Intel Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Intel and has not been reviewed or approved by Intel.

  • Leave & Time Off Breadth Sabbaticals and paid time off are highlighted as signature elements, with an established program offering four weeks after four years or eight weeks after seven years. This distinctive time off is positioned as a meaningful part of the overall package.
  • Parental & Family Support Paid bonding leave of 12 weeks and a New Parent Reintegration program, plus fertility benefits around $40,000 and up to $15,000 adoption reimbursement with no lifetime cap, are clearly stated. These programs are presented as standout components alongside broader family support resources.
  • Healthcare Strength Multiple medical plan options with 2026 updates, a shift to Spring Health for EAP, and in‑network virtual medical visits covered at 100% beginning in 2026 indicate a comprehensive offering. These features signal an emphasis on robust medical access and mental health support.

Intel Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Santa Clara, CA
75,000 Employees
Year Founded: 1968

What We Do

Our mission is to shape the future of technology to help create a better future for the entire world, that’s the power of Intel Inside. With more ingenuity and creativity inside, our work is at the heart of countless innovations. From major breakthroughs to things that make everyday life better— they’re all powered by Intel technology. With a career at Intel, you can help make the future more wonderful for everyone.

Similar Jobs

Magna International Logo Magna International

Senior Perception Engineer

Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Hybrid
Changning, Shanghai, CHN
171000 Employees

Mastercard Logo Mastercard

Manager, Products and Solutions - Agent Pay & Tokenized eCommerce Solutions

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Shanghai, Shanghai Municipality, Shanghai, CHN
38800 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sr. Manager, Marketplace Planning & Data Intelligence

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Shanghai, Shanghai Municipality, Shanghai, CHN
16000 Employees

Adyen Logo Adyen

Implementation Engineer I

Fintech • Payments • Financial Services
Easy Apply
Hybrid
Shanghai, Shanghai Municipality, Shanghai, CHN
4771 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account