SSE - Optimization Engineer

Posted 13 Days Ago
Be an Early Applicant
Ramapuram, Chennai, Tamil Nadu, IND
In-Office
Senior level
Information Technology • Consulting
The Role
Develop and optimize deep learning models for efficient inference across CPUs, GPUs, accelerators, and edge devices. Implement quantization algorithms, model compression, and high-performance C++ kernels. Optimize transformer, LLM, MoE, and PEFT workloads for latency, throughput, and memory usage. Profile and benchmark models across hardware platforms, translate research into production solutions, and collaborate with machine learning, compiler, and hardware teams.
Summary Generated by Built In
Title - Senior Software Engineer
 Job Description
Develop and optimize deep learning models (CNNs, LLMs, MoE) for efficient inference across CPU, GPU, and hardware accelerators / edge devices.
• Design and implement Quantization algorithms (PTQ, QAT, GPTQ, AWQ) from scratch.
• Apply model compression techniques such as pruning, decomposition, and distillation.
• Implement and optimize quantized kernels (INT8, INT4, FP8) using C++ for high performance.
• Translate research papers into production-ready implementations.
• Optimize latency, throughput, and memory usage for real-world deployment.
• Work on transformer optimization including KV-cache, PEFT (LoRA/QLoRA), and MoE models.
• Profile, benchmark, and debug model performance across different hardware platforms.
• Collaborate with ML, compiler, and hardware teams to deliver optimized solutions. 
Must-Have
• BE/BTech/MS/MTech in Computer Science or related field with 4+ years of experience.
• Strong programming skills in Python and C++.
• Proven experience in Quantization algorithms (PTQ, QAT, GPTQ, AWQ).
• Hands-on experience in pruning, model compression, and inference optimization.
• Experience implementing quantization or optimization techniques from scratch.
• Strong understanding of CNNs, Transformers, and LLM architectures.
• Experience with PyTorch / ONNX and model deployment pipelines.
• Strong problem-solving and performance optimization skills.
Nice-to-Have
• Experience with MoE architectures, and PEFT techniques (LoRA, QLoRA).
• Knowledge of TensorRT, ONNX Runtime, TVM, MLIR.
• Familiarity with hardware-aware optimization (GPU, NPU, Edge Devices).
• Experience in research paper implementation or open-source contributions.


Skills Required

  • BE, BTech, MS, or MTech in Computer Science or a related field
  • 4 or more years of professional experience
  • Strong programming skills in Python and C++
  • Proven experience with PTQ, QAT, GPTQ, and AWQ quantization algorithms
  • Hands-on experience with pruning, model compression, and inference optimization
  • Experience implementing quantization or optimization techniques from scratch
  • Strong understanding of CNN, Transformer, and LLM architectures
  • Experience with PyTorch, ONNX, and model deployment pipelines
  • Strong problem-solving and performance optimization skills
  • Experience with MoE architectures and PEFT techniques such as LoRA and QLoRA
  • Knowledge of TensorRT, ONNX Runtime, TVM, and MLIR
  • Familiarity with hardware-aware optimization for GPUs, NPUs, and edge devices
  • Experience implementing research papers or contributing to open source
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Chennai, TamilNadu
487 Employees
Year Founded: 2009

What We Do

MulticoreWare delivers software IP Solutions and Engineering Services serving a wide group of customers with Compilers & Toolchains, Libraries for SDK, Video codec and AI analytics solutions using various vision & non-vision (Radar, LiDAR, IMU, GPS, etc.) sensors on various heterogenous computing platforms. Our solutions are used in Automotive (ADAS/AD), Surveillance, Defence, Medical Imaging, IoT, Retail, Logistics, Industrial, Robotics, Smart City. MulticoreWare’s industry-leading video codec products (x266™/x265/Ultraziq) have been deployed in live streaming or VOD services across many broadcast customers. FOLLOW US ON SOCIAL MEDIA Youtube: https://www.youtube.com/channel/Multicoreware Twitter: https://twitter.com/MulticoreWare Facebook: https://www.facebook.com/multicoreware Instagram: https://www.instagram.com/multicoreware.inc/

Similar Jobs

Optum Logo Optum

Senior Manager Software Engineering

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chennai, Tamil Nadu, IND
160000 Employees

Optum Logo Optum

Manager AI/ML Engineering - Amazon Lex, Google ADK or Dialogflow

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chennai, Tamil Nadu, IND
160000 Employees

Optum Logo Optum

Senior Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Chennai, Tamil Nadu, IND
160000 Employees

Toast Logo Toast

Customer Care Specialist - xtraCHEF

Cloud • Fintech • Food • Information Technology • Software • Hospitality
In-Office
Chennai, Tamil Nadu, IND
5000 Employees

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account