Founding GPU Engineer

Reposted One Month Ago
Be an Early Applicant
2 Locations
In-Office or Remote
Mid level
Renewable Energy
The Role
Develop and optimize CUDA kernels and GPU-accelerated software for latency-sensitive, high-throughput workloads. Profile and tune GPU performance, scale multi-GPU/multi-node systems, build tooling to correlate GPU utilization with real-time energy/grid signals, collaborate with infrastructure and ML teams, and contribute to libraries and best practices for GPU performance in data center environments.
Summary Generated by Built In

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We've raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We're building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI, optimising how power-dense GPU workloads are scheduled, cooled and balanced against grid conditions in real time. We're looking for a Founding GPU Engineer to develop and optimise GPU-accelerated software for data centre systems: low-level performance engineering for large-scale compute clusters, tying GPU workload behaviour to energy availability and grid demand. This puts CUDA/GPU performance engineering at the centre of how Fuse scales its compute infrastructure.

Responsibilities
  • Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads
  • Profile and tune GPU performance across compute, memory bandwidth, and interconnect (NVLink/PCIe) bottlenecks
  • Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals
  • Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar communication libraries
  • Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity
  • Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines
  • Benchmark against CPU/GPU baselines and drive continuous performance improvements
  • Contribute to internal libraries, documentation, and best practices for GPU performance engineering

Requirements
  • 4+ years writing production CUDA code, or equivalent strong project/industry experience
  • Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy)
  • Proficiency in C++ and CUDA; experience with Python for tooling/orchestration
  • Experience with performance profiling tools (Nsight Systems/Compute)
  • Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand)
  • Strong grasp of memory optimisation, kernel fusion and parallel algorithm design
  • Comfortable working across the stack, from low-level kernels to system-level infrastructure
  • Bonus: Triton, cuDNN, cuBLAS or custom ML inference/training frameworks; data centre power/thermal management or demand-response systems; HPC, quantitative finance or large-scale distributed systems; Kubernetes/Slurm for GPU cluster orchestration; interest in energy markets, grid systems or sustainability-focused compute

Benefits
  • Competitive salary and eligibility for equity
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Private health insurance
  • Breakfast and dinner allowance for office-based employees

As we hire globally, benefits vary by location.

Skills Required

  • 4+ years of experience writing production CUDA code or equivalent
  • Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy)
  • Proficiency in C++ and CUDA
  • Experience with Python for tooling and orchestration
  • Experience with performance profiling tools (Nsight Systems/Compute)
  • Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand)
  • Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design
  • Comfortable working across the stack from low-level kernels to system-level infrastructure
  • Experience with Triton, cuDNN, cuBLAS, or custom ML inference/training frameworks
  • Exposure to data center power/thermal management or demand-response systems
  • Background in HPC, quantitative finance, or large-scale distributed systems
  • Familiarity with Kubernetes or Slurm for GPU cluster orchestration
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Greenwood Village, CO
97 Employees
Year Founded: 2022

What We Do

We are building a full stack renewable energy company to accelerate global renewable energy transition. We are based in London, New York. We've raised $90m from top tier investors like Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, MultiCoin. Energy is broken – We are here to fix it.

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
13 Locations
40000 Employees
45K-85K Annually

Boeing Logo Boeing

Trade Compliance Specialist 4 - Remote

Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
In-Office or Remote
Bingen, WA, USA
170000 Employees
99K-135K Annually
Remote or Hybrid
Chatsworth Lake Manor, CA, USA
205000 Employees
37K-66K Hourly

Comcast Logo Comcast

Senior Measurement & Attribution Analyst - Comcast Advertising

Digital Media • Information Technology • News + Entertainment
Remote or Hybrid
Virginia, USA
115000 Employees
78K-117K Annually

Similar Companies Hiring

UL Solutions Thumbnail
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Chicago, IL
15000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account