Acceleration Kernel Developer Lead

Reposted 23 Hours Ago
Be an Early Applicant
2 Locations
In-Office
Mid level
Hardware • Manufacturing
The Role
Design, implement, and optimize performance-critical GPU-style kernels (e.g., matmul, attention) and host-side orchestration. Identify bottlenecks, deliver measurable throughput improvements, create micro-benchmarks, regression tests, and tooling, and collaborate with compiler, runtime, ML, and hardware teams to integrate kernels into production.
Summary Generated by Built In

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.

As the Acceleration Kernel Developer Lead at Tenstorrent, you will take on a pivotal role in guiding the optimization of low-level workloads, kernel development, and enhancing the performance of our software for machine learning applications. You will lead a team of highly skilled engineers, ensuring our software operates at peak efficiency and delivers high-quality results to our clients and users.

This role is hybrid based out of Warsaw or Gdansk, Poland.

We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.


Who You Are

  • An experienced technical leader who thrives on guiding talented engineers through complex, low-level software challenges.
  • A performance-obsessed optimizer with an analytical mindset for diagnosing hardware-software bottlenecks.
  • A collaborative communicator who effortlessly translates machine learning requirements into highly efficient low-level code.
  • A self-motivated innovator who stays at the absolute forefront of machine learning compilation and high-performance computing (HPC) trends.

What We Need

  • Strong proficiency in C/C++ with a proven track record of writing high-performance low-level code.
  • Hands-on experience developing and optimizing tensor compute or data movement kernels.
  • Deep familiarity with performance profiling tools and debugging complex runtime environments.
  • Solid understanding of machine learning frameworks, with bonus points for GPU programming (CUDA, OpenCL) or operating system internals.

What You Will Learn

  • How to program and unlock the full potential of Tenstorrent's unique, networked RISC-V and AI processor architectures.
  • Advanced hardware-software co-design techniques that unify machine learning compilation with custom silicon execution.
  • To scale next-generation LLMs and HPC workloads across massive, highly parallel compute grids.

Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.

This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology.  Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2).   These requirements apply to persons located in the U.S. and all countries outside the U.S.  As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency.  If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.

Skills Required

  • Strong C++ systems engineering experience
  • Experience reasoning about concurrency, synchronization, latency hiding, and compute vs memory trade-offs
  • Data-driven optimization using profiling and benchmarking
  • Debugging complex runtime or kernel-level issues in large codebases
  • Designing, implementing, and optimizing GPU-style kernels (matrix multiplication, attention primitives, data-movement ops)
  • Ownership of performance tuning from bottleneck identification to throughput improvements
  • Contribution to host-side orchestration code and parallelization strategies
  • Development of micro-benchmarks, regression tests, and tooling for correctness and performance
  • Close collaboration with compiler, runtime, ML, and hardware teams
  • Hybrid work based out of Warsaw or Gdansk, Poland
  • Eligibility to access U.S. export-controlled technology (citizenship/permanent residency or ability to obtain license approval)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Toronto, ON
389 Employees
Year Founded: 2016

What We Do

Tenstorrent is a next-generation computing company that builds computers for AI. Headquartered in Toronto, Canada, with U.S. offices in Austin, Texas, and Silicon Valley, and global offices in Belgrade and Bangalore, Tenstorrent brings together experts in the field of computer architecture, ASIC design, advanced systems, and neural network compilers. Join us: www.tenstorrent.com/careers

Similar Jobs

Capco Logo Capco

Test Automation Engineer

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Capco Logo Capco

Manual Tester – Cards (Polish is Mandatory)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Pfizer Logo Pfizer

Vice President, Strategy, Value & Innovation

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
43 Locations
121990 Employees

Pfizer Logo Pfizer

Vice President, Build - Data, Engineering & AI

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
43 Locations
121990 Employees

Similar Companies Hiring

Fortune Brands Innovations Thumbnail
Manufacturing
Deerfield, IL
10000 Employees
Fairly Even Thumbnail
Hardware • Robotics • Sales • Software • Hospitality
New York, NY
30 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account