Lead the integration of diverse AI models including VLA, Vision, and Multimodal architectures by utilizing our kernel programming language to ensure both accuracy and performance while keeping the stack ready for developers to use.
Responsibilities
Design and implement efficient kernels on FuriosaAI’s kernel programming stack (including vISA, TCL), targeting Tensor Contract Processor (TCP) architectures.
Diagnose and optimize kernel performance with profiling tools and roofline analysis for each RNGD-accelerated AI model.
Develop and apply automated kernel generation and optimization for AI workloads.
Build diagnostic tools or testbeds for robust and reliable kernel validation.
Drive end-to-end programming enablement on RNGDs, creating reproducible guides and reference implementations.
BS in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.
Experience in low-level systems programming targeting XPU (e.g., NPU, GPU) architectures.
Experience collaborating across engineering, research, and product teams to align software development with product requirements.
MS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.
Experience in optimizing high-performance kernels on AI accelerators (e.g., GPU, TPU) for AI products.
Understanding of XPU architecture (computation patterns, data movement) and software-hardware co-optimization strategies.
Experience in open-source or research projects on AI model architectures such as Diffusion, Mamba, and VLA.
Experience in designing efficient deep learning architectures and developing algorithms for AI applications.
Skills Required
- BS in Computer Science, Artificial Intelligence, Electrical Engineering, or related field
- Experience in low-level systems programming targeting XPU (e.g., NPU, GPU) architectures
- Experience collaborating across engineering, research, and product teams
- MS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or related field
- Experience optimizing high-performance kernels on AI accelerators (e.g., GPU, TPU)
- Understanding of XPU architecture and software-hardware co-optimization strategies
- Experience in open-source or research projects on AI model architectures (Diffusion, Mamba, VLA)
- Experience designing efficient deep learning architectures and developing algorithms for AI applications
What We Do
FuriosaAI designs and develops data center accelerators for the most advanced AI models and applications. Our mission is to make AI computing sustainable so everyone on Earth has access to powerful AI. Our Background Three misfit engineers with each from HW, SW and algorithm fields who had previously worked for AMD, Qualcomm and Samsung got together and founded FuriosaAI in 2017 to build the world’s best AI chips. The company has raised more than $100 million, with investments from DSC Investment, Korea Development Bank, and Naver, the largest internet provider in Korea. We have partnered on our first two products with a wide range of industry leaders including TSMC, ASUS, SK Hynix, GUC, and Samsung. FuriosaAI now has over 140 employees across Seoul, Silicon Valley, and Europe. Our Approach We are building full stack solutions to offer the most optimal combination of programmability, efficiency, and ease of use. We achieve this through a “first principles” approach to engineering: We start with the core problem, which is how to accelerate.









