Location : Bengaluru
Key Responsibilities :
• Develop, optimize, and maintain GPU-accelerated components for deep learning pipelines using frameworks such as CUDA, HIP, or OpenCL
• Analyse and improve GPU kernel performance through profiling, benchmarking, and resource optimization.
• Optimize memory access, compute, throughput, and kernel execution to improve overall system performance on the target GPUs.
• Port existing CPU-based implementations to GPU platforms while ensuring correctness and performance scalability.
• Work closely with system architects, software engineers, and domain experts to integrate GPU-accelerated solutions.
Required Qualifications :
• Bachelor's or master's degree in computer science, Electrical Engineering, or a related field.
• 2+ years of hands-on experience in GPU programming, preferably using CUDA, or other GPU APIs like HIP, OpenCL etc.,
• Strong understanding of GPU architecture, memory hierarchy, shared memory, bank conflicts and parallel programming models.
• Proficiency in C/C++ and hands-on experience developing on Linux-based systems.
• Familiarity with profiling and tuning tools such as Nsight, rocprof, or Perfetto
Good to have skills in addition to GPU :
• Knowledge and Experience in SIMD Programming
• Good understanding of NN Operators & Hands-on experience with PT, TF, Tensor RT.
• Exposure to DL Concepts like Quantization, Pruning etc., • Experience in working with High Performance Compute (HPC) Systems
Skills Required
- Bachelor's or master's degree in computer science, electrical engineering, or a related field
- 2+ years of hands-on experience in GPU programming
- Experience with CUDA, HIP, OpenCL, or other GPU programming APIs
- Strong understanding of GPU architecture, memory hierarchy, shared memory, bank conflicts, and parallel programming models
- Proficiency in C/C++
- Hands-on experience developing on Linux-based systems
- Familiarity with Nsight, rocprof, or Perfetto profiling and tuning tools
- Knowledge and experience in SIMD programming
- Understanding of neural-network operators and hands-on experience with PyTorch, TensorFlow, or TensorRT
- Exposure to deep learning concepts such as quantization and pruning
- Experience with high-performance computing systems
What We Do
MulticoreWare delivers software IP Solutions and Engineering Services serving a wide group of customers with Compilers & Toolchains, Libraries for SDK, Video codec and AI analytics solutions using various vision & non-vision (Radar, LiDAR, IMU, GPS, etc.) sensors on various heterogenous computing platforms. Our solutions are used in Automotive (ADAS/AD), Surveillance, Defence, Medical Imaging, IoT, Retail, Logistics, Industrial, Robotics, Smart City. MulticoreWare’s industry-leading video codec products (x266™/x265/Ultraziq) have been deployed in live streaming or VOD services across many broadcast customers. FOLLOW US ON SOCIAL MEDIA Youtube: https://www.youtube.com/channel/Multicoreware Twitter: https://twitter.com/MulticoreWare Facebook: https://www.facebook.com/multicoreware Instagram: https://www.instagram.com/multicoreware.inc/







