System Software Engineer - AI

Posted 22 Days Ago
Hiring Remotely in USA
Remote
140K-200K Annually
Mid level
Artificial Intelligence • Hardware • Semiconductor
Stronger Network. Faster Inference.
The Role
As a System Software Engineer, you will design communication primitives for AI models, optimize performance for training workloads, and benchmark performance on clusters.
Summary Generated by Built In

System Software Engineer - AI

About us:

We are a stealth-mode startup building foundational technology to address performance, scalability, and resiliency challenges in large-scale AI data center clusters. We are backed by top-tier VC firms and notable angel investors.

The company is led by experienced builders and operators who have founded companies, taken them to scale, and exited successfully. We work with a strong sense of unity and shared responsibility, and we expect trust, integrity, and respect in how we collaborate and make decisions. We hold ourselves accountable to one another and to the quality of the work we deliver.

Headquartered in Silicon Valley, we operate across a mix of remote and on-site locations in the U.S. and Canada. We aim to create an environment where people are treated fairly, supported in their growth, and are empowered to do meaningful work alongside others who take the craft seriously.

We are looking for:

We are looking for a talented System Software Engineer to help us redefine the infrastructure layer of AI. In this role, you will bridge the gap between high-level AI frameworks and low-level system software. You will be responsible for designing and implementing the communication and execution primitives that allow large-scale AI models to run efficiently across thousands of GPUs. We are looking for a "builder" who thrives in the early stages of a product’s lifecycle and is passionate about solving the "hard" systems problems of the generative AI era.

Key Responsibilities:

  • Collaborate across the stack to influence the design of our foundational technology, ensuring it meets the needs of next-generation AI models.

  • Identify and resolve performance bottlenecks in distributed training and inference workloads through deep-dive analysis of the software-hardware interface.

  • Conduct rigorous performance benchmarking and characterization on multi-node clusters.

Required Skills and Qualifications:

  • Strong proficiency in C++ and Python, with a deep understanding of systems programming fundamentals (memory management, concurrency, OS internals).

  • Proficient in a Linux development environment.

Desired Skills:

  • Experience with GPU programming (CUDA) and performance optimization for parallel architectures.

  • Familiarity with distributed AI frameworks (PyTorch, JAX, or DeepSpeed) and/or inference engines (vLLM, SGLang, Dynamo/TRT-LLM).

  • Hands-on experience with large-scale cluster orchestration and telemetry tools.

Education:

  • Bachelor's or Master's degree in Computer Engineering, Computer Science, or a related field.

Compensation:

Target base salary for this role is $140,000 - $200,000 per year + meaningful equity + benefits + 401k. Our salary ranges are determined by role, level, experience, and location.

We are an equal opportunity employer. We value a range of perspectives and experiences and make employment decisions based on merit and business needs. We do not discriminate on the basis of legally protected characteristics.

Agency Note:

We do not accept resumes from agencies or search firms. Please do not forward candidate profiles through our careers page, email, LinkedIn messages, or directly to company employees. Any resumes submitted will be deemed the property of the company, and no fees will be paid in the event the candidate is hired.

#LI-EW1

Top Skills

C++
Cuda
Deepspeed
Dynamo/Trt-Llm
Jax
Linux
Python
PyTorch
Sglang
Vllm
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
40 Employees
Year Founded: 2025

What We Do

We are a mixture of software, system, and silicon experts using AI every day to deliver the world's most capable and responsive intelligence.

Why Work With Us

We are a mixture of software, system, and silicon experts using AI every day to deliver the world's most capable and responsive intelligence. We start from the workload. Scaling inference is less about brute force and more about how compute is distributed and memory is interc

Similar Jobs

Golden Hippo Logo Golden Hippo

Paid Ads Specialist

Digital Media • eCommerce • Information Technology • Marketing Tech • Retail • Social Media • Analytics
Remote or Hybrid
2 Locations
500 Employees
72K-90K Annually

Golden Hippo Logo Golden Hippo

Sales Manager

Digital Media • eCommerce • Information Technology • Marketing Tech • Retail • Social Media • Analytics
Remote or Hybrid
Northern California, CA, USA
500 Employees
78K-104K Annually

Zscaler Logo Zscaler

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
California, USA
8697 Employees
90K-129K Annually

Dandy Logo Dandy

Staff Software Engineer

Computer Vision • Healthtech • Information Technology • Logistics • Machine Learning • Software • Manufacturing
Remote
USA
1800 Employees
228K-285K Annually

Similar Companies Hiring

Idler Thumbnail
Artificial Intelligence
San Francisco, California
6 Employees
Fairly Even Thumbnail
Software • Sales • Robotics • Other • Hospitality • Hardware
New York, NY
Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account