Senior Researcher - Edge AI Optimization/Hardware-Aware ML

Posted 7 Days Ago
Be an Early Applicant
Edmonton, AB, CAN
In-Office
Senior level
Information Technology • Other
The Role
Research and develop hardware-aware ML optimization techniques for edge devices, including quantization, pruning, efficient architectures, compiler and runtime optimization, and heterogeneous CPU/GPU/NPU deployment. Profile real-device performance, energy, memory, and latency; build benchmarking and regression infrastructure; optimize LLM/VLM inference; publish findings and patents; collaborate with product, platform, and hardware teams; and mentor junior researchers while shaping the technical roadmap.
Summary Generated by Built In

Huawei Canada has an immediate permanent opening for a Researcher.

About the team:

The Software-Hardware System Optimization Lab focuses on research and innovation in power efficiency and performance optimization for consumer devices. By leveraging the talents and capabilities of local academia and our team, we aim to build system-optimization capabilities for software and hardware across edge AI, multimedia, graphics, mobile gaming, and system software domains, thereby enhancing the user experience and performance competitiveness of Huawei's consumer device products.


About the job:

  • Conduct research in hardware-aware neural network optimization (e.g., quantization-aware training, mixed precision, pruning, distillation, neural architecture search).

  • Develop novel approaches for latency/energy-aware training objectives and Pareto optimization (accuracy vs. compute vs. memory).

  • Prototype and evaluate techniques for efficient inference under device constraints (thermal limits, memory bandwidth, intermittent connectivity).

  • Publish and present findings internally and externally (papers, workshops, patents, technical blogs).

  • Optimize inference pipelines across pre/post-processing, scheduling, operator fusion, memory planning, and runtime execution.

  • Collaborate on or contribute to compilers / runtimes (e.g., TVM, MLIR, XLA, TensorRT, ONNX Runtime, TFLite, ExecuTorch) to improve operator coverage and performance.

  • Profile and optimize models with real device traces, addressing bottlenecks such as cache misses, memory bandwidth, kernel launch overhead, and CPU–NPU handoff.

  • Build and maintain hardware-aware benchmarking methodology and regression suites for edge targets (ARM CPU, mobile GPU, DSP, NPU).

  • Create deployment recipes for heterogeneous compute (CPU+GPU+NPU) including partitioning strategies and fallback paths.

  • Drive optimization for on-device personalization and incremental updates when needed (e.g., small adapters, efficient fine-tuning).

  • Partner with product engineering, platform teams, and hardware teams to translate device constraints into research targets and to transition research prototypes into production.

  • Mentor junior researchers/engineers, review experimental designs, and raise the quality bar for measurement rigor and reproducibility.

  • Define technical roadmap areas (e.g., next-gen quantization, kernel optimization, model families for edge, compiler improvements).

About the ideal candidate:

  • PhD (or equivalent research experience) in Machine Learning, Computer Science, Electrical/Computer Engineering, or related field.

  • Strong programming skills in Python and C/C++ (or equivalent systems language).Experience building AI agent / harness / skill toolchains, including model evaluation, orchestration, and LLM-powered tooling. Hands-on experience with deep learning frameworks (e.g., PyTorch, TensorFlow, JAX) and deployment toolchains (e.g., ONNX, TFLite, TensorRT, TVM, MLIR-based stacks). Solid knowledge of performance profiling: latency measurement, memory profiling, kernel-level bottleneck analysis, and experimental rigor.

  • Proven publication record at top venues (e.g., NeurIPS/ICML/ICLR, MLSys, ASPLOS, ISCA, MICRO) and/or patents in ML efficiency.

  • 2+ years of relevant experience (research lab or industry) with demonstrated impact in at least one of:

    • Model compression (quantization/pruning/distillation)

    • Efficient architectures (MobileNet-like, MoE, efficient transformers, etc.)

    • ML systems/compilers/runtime optimization

    • hardware-aware optimization for edge deployment

  • Experience optimizing for specific edge hardware:

    • ARM NEON, mobile GPUs, DSPs, NPUs, microcontrollers

    • Experience with distributed benchmarking, CI for performance regression, and reproducible experiment pipelines.

    • Understanding of power/thermal constraints and methodologies for measuring energy on device.

    • Experience with efficient LLM/VLM inference on edge (KV-cache optimization, quantized attention, speculative decoding, etc.).

  • Technical Skills:

    • Quantization: PTQ/QAT, per-channel/per-tensor, calibration, smooth quant, GPTQ-like methods, mixed precision

    • Sparsity: structured pruning, N: M sparsity, hardware-friendly sparsity

    • Compiler techniques: graph rewriting, operator lowering, scheduling, kernel autotuning

    • Runtime techniques: memory arenas, tensor lifetime analysis, static vs dynamic shapes, batching strategies

    • Hardware fundamentals: cache hierarchy, SIMD, memory bandwidth, accelerator programming models

Additional Information:

Huawei Canada is committed to a fair, inclusive, and accessible recruitment process. If you require accommodation during any stage of the hiring process, please let us know and we will work with you to meet your needs.

All applications for this position are reviewed directly by our hiring team, we do not use artificial intelligence tools to screen or select candidates.

Skills Required

  • PhD or equivalent research experience in Machine Learning, Computer Science, Electrical/Computer Engineering, or a related field
  • Strong programming skills in Python and C/C++ or an equivalent systems language
  • Experience building AI agent, harness, or skill toolchains involving model evaluation, orchestration, and LLM-powered tooling
  • Hands-on experience with deep learning frameworks such as PyTorch, TensorFlow, or JAX
  • Experience with deployment toolchains such as ONNX, TFLite, TensorRT, TVM, or MLIR-based stacks
  • Knowledge of latency measurement, memory profiling, kernel-level bottleneck analysis, and rigorous experimentation
  • Publication record at top venues and/or patents in ML efficiency
  • At least 2 years of relevant research laboratory or industry experience
  • Demonstrated impact in model compression, efficient architectures, ML systems/compiler/runtime optimization, or hardware-aware edge deployment
  • Experience optimizing for ARM NEON, mobile GPUs, DSPs, NPUs, or microcontrollers
  • Experience with distributed benchmarking, performance-regression CI, and reproducible experiment pipelines
  • Understanding of power and thermal constraints and on-device energy measurement methodologies
  • Experience with efficient edge LLM/VLM inference, including KV-cache optimization, quantized attention, or speculative decoding
  • Knowledge of quantization, sparsity, compiler techniques, runtime techniques, cache hierarchy, SIMD, memory bandwidth, and accelerator programming models
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Shenzhen
1,770 Employees
Year Founded: 1987

What We Do

Founded in 1987, Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices. We are committed to bringing digital to every person, home and organization for a fully connected, intelligent world. We have approximately 197,000 employees and we operate in over 170 countries and regions, serving more than three billion people around the world. In Canada, Huawei conducts innovative and leading edge research in 5G technologies, along with advanced development of emerging cloud, device and network technologies & services. While our renowned Canada Research Centre in the thriving technology landscape of Ottawa, Ontario continues to grow rapidly in size and strategic product initiatives, additional presence has also been established across Canada with R&D facilities in Vancouver, Edmonton, Waterloo, Markham, Montreal, and a R&D office in Quebec City.

Similar Jobs

Pfizer Logo Pfizer

Director R&D EHS Program Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees
177K-294K Annually
Remote or Hybrid
9 Locations
2449 Employees
86K-127K Annually

Block Logo Block

Machine Learning Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
277K-415K Annually

Block Logo Block

Account Executive

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
97K-172K Annually

Similar Companies Hiring

Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees
OmniCable Thumbnail
Other
Houston, Texas
815 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account