Performance Tools Intern

Posted 10 Days Ago
Be an Early Applicant
San Jose, CA, USA
In-Office
Internship
Artificial Intelligence • Hardware • Software
The Role
Design and build performance analysis and profiling tooling for a custom PCIe ML accelerator. Collect hardware counters and traces, correlate events across CPU, accelerator, storage and networking, and create visualization and analysis tools. Collaborate with hardware, compiler, firmware, and inference engineers to identify bottlenecks and improve performance.
Summary Generated by Built In

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

Join our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.

During your internship, you may:

  • Build components of our performance analysis and profiling infrastructure.

  • Collect and analyze performance data from our custom ML accelerators, including hardware counters, execution traces, and memory behavior.

  • Develop tooling to trace host-side runtime activity, system behavior, and accelerator execution.

  • Help correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads.

  • Build analysis and visualization tools that help engineers identify performance bottlenecks and optimize models.

  • Work alongside hardware, compiler, firmware, and inference engineers to understand performance challenges and develop tools that improve developer productivity.

Representative projects

  • Implement the data collection framework for hardware performance counters on a custom PCIe-based accelerator.

  • Develop a user-space service for low-overhead tracing of accelerator activity.

  • Design and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.

  • Create an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.

You may be a good fit if you have

  • Strong programming skills in C++ or Rust. Experience with Python is a plus.

  • Solid understanding of computer architecture, including CPUs, GPUs or AI accelerators, memory hierarchies, and parallel programming.

  • Experience or strong interest in low-level performance analysis, profiling, and performance optimization.

  • Familiarity with performance analysis tools such as Nsight, VTune, Xprof, Perfetto, or similar tools is a plus.

  • Experience or strong interest in operating systems, compilers, firmware, drivers, or other low-level systems software.

  • Passion for understanding how complex systems behave under real workloads and building tools that help other engineers optimize performance.

  • Strong problem-solving skills and curiosity to learn quickly in a fast-paced engineering environment.

Strong candidates may also have experience with (Nice-to-have qualifications)

  • Direct experience developing performance analysis or debugging tools.

  • Experience with ML accelerator architectures (GPUs, TPUs, etc.).

  • Experience with kernel-mode driver development (Linux or Windows).

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

 

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Skills Required

  • Strong programming skills in C++ or Rust
  • Experience with Python
  • Solid understanding of computer architecture (CPUs, GPUs/accelerators, memory hierarchies)
  • Experience or strong interest in low-level performance analysis, profiling, and optimization
  • Familiarity with performance analysis tools such as Nsight, VTune, Xprof, Perfetto
  • Experience or strong interest in operating systems, compilers, firmware, drivers, or low-level systems software
  • Direct experience developing performance analysis or debugging tools
  • Experience with ML accelerator architectures (GPUs, TPUs)
  • Experience with kernel-mode driver development (Linux or Windows)

Etched Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Etched and has not been reviewed or approved by Etched.

  • Equity Value & Accessibility Equity growth is described as strong and significant equity is part of the package. High total compensation for technical roles reinforces the equity-led upside.
  • Healthcare Strength Medical, dental, and vision coverage include generous premium support, indicating robust core healthcare. This reduces employee cost exposure for essential coverage.
  • Wellbeing & Lifestyle Benefits Daily lunch and dinner, a housing subsidy for those living near the office, relocation support, and wellness perks are highlighted. These offerings lower day-to-day living costs and support practical wellbeing.

Etched Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Pakenham
53 Employees
Year Founded: 2022

What We Do

By burning the transformer architecture into our chips, we’re creating the world’s most powerful servers for transformer inference.

Similar Jobs

Shield AI Logo Shield AI

Cybersecurity Engineer

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
In-Office
San Diego, CA, USA
160K-240K Annually

ServiceNow Logo ServiceNow

Sr Data Analytics Mgr. Paid Media Campaigns

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
San Diego, CA, USA
29000 Employees
140K-245K Annually

ServiceNow Logo ServiceNow

Senior Customer Success Manager

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
San Diego, CA, USA
29000 Employees

Motive Logo Motive

Director, Developer Platform & Experience

Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation
Easy Apply
In-Office
2 Locations
4000 Employees
229K-285K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account