Sr. Software Engineer - Inference Engine (Platform Software)

Posted 3 Days Ago
Be an Early Applicant
Seoul, KOR
In-Office
Senior level
Artificial Intelligence • Information Technology • Software • Database • Manufacturing
The Role
Design, implement, and optimize a high-performance inference engine for LLMs on FuriosaAI NPUs. Implement advanced inference optimizations (speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution), enable distributed/scalable inference, collaborate with compiler and hardware teams, and integrate state-of-the-art LLM serving techniques into production.
Summary Generated by Built In
About the Job

Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.

In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential.

Responsibilities
  • Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency.

  • Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine.

  • Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake.

  • Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.

  • Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine.

Minimum Qualifications
  • BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience

  • Proficiency in Rust or C++ programming skill

  • Knowledge and passion of deep learning, LLM, and/or generative AI models

  • Excellent problem-solving and data analysis skills.

  • Strong communication and collaboration skills.

Preferred Qualifications
  • Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing.

  • A deep understanding of performance optimization systems.

  • Proficiency in C++/CUDA or Triton kernel development

  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM.

Contact

Skills Required

  • BS degree in Computer Science, Engineering, or related field, with at least 3 years of relevant industry experience, or equivalent practical experience
  • Proficiency in Rust or C++ programming skill
  • Knowledge and passion of deep learning, LLM, and/or generative AI models
  • Excellent problem-solving and data analysis skills
  • Strong communication and collaboration skills
  • Experience in building inference serving systems for large models, including batching, scheduling, caching, and load balancing
  • A deep understanding of performance optimization systems
  • Proficiency in C++/CUDA or Triton kernel development
  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Seoul, Seoul
143 Employees
Year Founded: 2017

What We Do

FuriosaAI designs and develops data center accelerators for the most advanced AI models and applications. Our mission is to make AI computing sustainable so everyone on Earth has access to powerful AI. Our Background Three misfit engineers with each from HW, SW and algorithm fields who had previously worked for AMD, Qualcomm and Samsung got together and founded FuriosaAI in 2017 to build the world’s best AI chips. The company has raised more than $100 million, with investments from DSC Investment, Korea Development Bank, and Naver, the largest internet provider in Korea. We have partnered on our first two products with a wide range of industry leaders including TSMC, ASUS, SK Hynix, GUC, and Samsung. FuriosaAI now has over 140 employees across Seoul, Silicon Valley, and Europe. Our Approach We are building full stack solutions to offer the most optimal combination of programmability, efficiency, and ease of use. We achieve this through a “first principles” approach to engineering: We start with the core problem, which is how to accelerate.

Similar Jobs

Legora Logo Legora

Senior Legal Engineer - South Korea

Artificial Intelligence • Legal Tech • Software
In-Office
Seoul, KOR
700 Employees

Legora Logo Legora

Engagement Manager

Artificial Intelligence • Legal Tech • Software
In-Office
Seoul, KOR
700 Employees

Airwallex Logo Airwallex

Manager, Revenue Strategy & Operations (Korea GTM)

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
In-Office
Seoul, KOR
2300 Employees

Airwallex Logo Airwallex

Senior Associate/Manager, Revenue Strategy & Operations (Korea GTM)

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
In-Office
Seoul, KOR
2300 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account