GPU Systems Engineer

Reposted 4 Days Ago
Be an Early Applicant
Bellevue, WA, USA
In-Office
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
The Role
Design and optimize GPU workloads for AI, high-performance computing, and data-processing systems. Develop CUDA kernels, custom ML operators, optimized libraries, benchmarks, and regression tests. Profile GPU code, tune memory and execution behavior, optimize distributed multi-GPU training, and evaluate new accelerator architectures. Collaborate with ML and engineering teams, contribute to compiler-level optimizations, document performance practices, and mentor junior engineers.
Summary Generated by Built In
GPU Systems Engineer - Remote 
 
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. 
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. 
 
Job Title: GPU Systems Engineer
Location: 100% Remote (U.S.) 
Position Type: Full-time, Direct W2 
Salary Range: $100,000–$150,000 Annually 
Experience Required: 6+ years 
 
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. 
 
Job Summary 
We are seeking a GPU Systems Engineer with deep expertise in CUDA programming, GPU architecture, and high-performance computing to design and optimize compute-intensive workloads on modern accelerator hardware. This role focuses on extracting maximum performance from GPU platforms for AI training, inference, scientific computing, and high-throughput data processing workloads. The ideal candidate combines low-level systems mastery with strong software engineering practices, and has a track record of delivering measurable performance improvements on production GPU systems. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. 
Key Responsibilities 
  • Design and implement high-performance CUDA kernels for compute-intensive workloads across AI and HPC use cases. 
  • Profile and optimize GPU code using tools such as Nsight Systems, Nsight Compute, and CUDA profilers. 
  • Tune memory access patterns, occupancy, register usage, and shared memory utilization for peak performance. 
  • Develop highly optimized libraries for linear algebra, attention, and other ML primitives. 
  • Optimize multi-GPU and multi-node training using NCCL, RDMA, and high-performance networking. 
  • Implement custom operators and fused kernels in PyTorch, JAX, or Triton. 
  • Collaborate with ML engineers to identify performance bottlenecks in training and inference pipelines. 
  • Develop benchmarks and regression tests to safeguard performance over time. 
  • Evaluate new GPU architectures and feature sets, and advise on adoption strategy. 
  • Contribute to compiler-level optimizations for tensor programs where appropriate, working at the boundary between ML frameworks and underlying accelerator codegen to unlock performance not reachable through framework-level tuning alone. 
  • Optimize memory hierarchy usage across HBM, L2, shared memory, and registers. 
  • Implement mixed-precision and quantized compute paths that maximize accelerator throughput while preserving numerical fidelity within bounds acceptable for the target workloads. 
  • Document performance characteristics, design decisions, and tuning playbooks for internal teams. 
  • Stay current with GPU architecture, CUDA evolution, and emerging accelerator technologies. 
Required Qualifications 
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field. 
  • Six or more years of experience in GPU programming and performance engineering. 
  • Deep expertise in CUDA C/C++ and GPU programming models. 
  • Strong understanding of modern GPU architectures, memory hierarchies, and execution models. 
  • Hands-on experience profiling and optimizing GPU workloads in production. 
  • Familiarity with NCCL, MPI, and high-performance interconnect technologies. 
  • Experience integrating custom kernels into ML frameworks. 
  • Strong C++ skills and familiarity with modern systems programming practices. 
  • Solid grounding in linear algebra and numerical methods. 
  • Strong communication and collaboration skills with research and engineering teams. 
Preferred Qualifications 
  • Experience with Triton, CUTLASS, or other GPU kernel authoring frameworks. 
  • Familiarity with TensorRT, FasterTransformer, or vLLM internals. 
  • Exposure to compiler infrastructure such as LLVM or MLIR. 
  • Open-source contributions to GPU or ML performance libraries. 
  • Experience with large-scale distributed training infrastructure. 
How to Apply 
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com. 
Bright Vision Technologies is an Equal Opportunity Employer. 
 

Skills Required

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field
  • Six or more years of experience in GPU programming and performance engineering
  • Deep expertise in CUDA C/C++ and GPU programming models
  • Strong understanding of modern GPU architectures, memory hierarchies, and execution models
  • Hands-on experience profiling and optimizing GPU workloads in production
  • Familiarity with NCCL, MPI, and high-performance interconnect technologies
  • Experience integrating custom kernels into machine learning frameworks
  • Strong C++ skills and familiarity with modern systems programming practices
  • Solid grounding in linear algebra and numerical methods
  • Strong communication and collaboration skills with research and engineering teams
  • Experience with Triton, CUTLASS, or other GPU kernel authoring frameworks
  • Familiarity with TensorRT, FasterTransformer, or vLLM internals
  • Exposure to compiler infrastructure such as LLVM or MLIR
  • Open-source contributions to GPU or machine learning performance libraries
  • Experience with large-scale distributed training infrastructure
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
53 Employees
Year Founded: 2020

What We Do

Bright Vision Technologies is a minority-owned organization founded in July 2020 and based in New Jersey, USA. The company specializes in delivering top-tier staffing and IT consulting services, including custom computer programming and systems design. Additionally, they are a product engineering firm with a flagship AI-powered talent intelligence and enterprise automation platform called Lumina, which helps transform IT into a strategic asset for their valued partners.

Similar Jobs

Bright Vision Technologies Logo Bright Vision Technologies

Systems Engineer

Artificial Intelligence • Information Technology • Software • Consulting
In-Office
Sammamish, WA, USA
53 Employees
100K-150K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Acting Supervisor III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
North Bend, WA, USA
16000 Employees
16-25 Hourly

Arm Logo Arm

User Experience Researcher

Artificial Intelligence • Internet of Things • Semiconductor
Hybrid
3 Locations
8314 Employees
212K-286K Annually

Samsara Logo Samsara

Analytics Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
119K-180K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account