• Experience with MoE architectures, and PEFT techniques (LoRA, QLoRA).
Skills Required
- BE, BTech, MS, or MTech in Computer Science or a related field
- 4 or more years of professional experience
- Strong programming skills in Python and C++
- Proven experience with PTQ, QAT, GPTQ, and AWQ quantization algorithms
- Hands-on experience with pruning, model compression, and inference optimization
- Experience implementing quantization or optimization techniques from scratch
- Strong understanding of CNN, Transformer, and LLM architectures
- Experience with PyTorch, ONNX, and model deployment pipelines
- Strong problem-solving and performance optimization skills
- Experience with MoE architectures and PEFT techniques such as LoRA and QLoRA
- Knowledge of TensorRT, ONNX Runtime, TVM, and MLIR
- Familiarity with hardware-aware optimization for GPUs, NPUs, and edge devices
- Experience implementing research papers or contributing to open source
What We Do
MulticoreWare delivers software IP Solutions and Engineering Services serving a wide group of customers with Compilers & Toolchains, Libraries for SDK, Video codec and AI analytics solutions using various vision & non-vision (Radar, LiDAR, IMU, GPS, etc.) sensors on various heterogenous computing platforms. Our solutions are used in Automotive (ADAS/AD), Surveillance, Defence, Medical Imaging, IoT, Retail, Logistics, Industrial, Robotics, Smart City. MulticoreWare’s industry-leading video codec products (x266™/x265/Ultraziq) have been deployed in live streaming or VOD services across many broadcast customers. FOLLOW US ON SOCIAL MEDIA Youtube: https://www.youtube.com/channel/Multicoreware Twitter: https://twitter.com/MulticoreWare Facebook: https://www.facebook.com/multicoreware Instagram: https://www.instagram.com/multicoreware.inc/







