Performance Engineer

Posted 10 Days Ago
Be an Early Applicant
San Jose, CA, USA
In-Office
135K-170K Annually
Senior level
Big Data • Information Technology
The Role
Define and measure scale-up fabric performance for GPU clusters by building roofline models, benchmarks, and end-to-end workload studies. Run and scale real inference workloads, identify bottlenecks, perform head-to-head comparisons, and build automated test infrastructure. Partner with cross-functional teams to influence architecture, firmware, and product decisions and produce customer-facing performance reports and documentation.
Summary Generated by Built In

Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the full potential of modern AI. Astera Labs’ Intelligent Connectivity Platform integrates CXL®, Ethernet, NVLink, PCIe®, and UALink™ semiconductor-based technologies with the company’s COSMOS software suite to unify diverse components into cohesive, flexible systems that deliver end-to-end scale-up, and scale-out connectivity. The company’s custom connectivity solutions business complements its standards-based portfolio, enabling customers to deploy tailored architectures to meet their unique infrastructure requirements. Discover more at www.asteralabs.com.


Senior Performance Engineer 

Location: San Jose, CA (On-site) 

Role Overview 

Astera Labs is a hyper-growth connectivity company enabling the rack-scale AI infrastructure powering the world's most advanced GPU clusters. Our Scorpio scale-up fabric switches are purpose-built to unlock the performance of next-generation AI workloads, and we're looking for a Senior Performance Engineer to demonstrate the real-world value of our silicon where it matters most: on real inference and training workloads running on GPUs at scale. 

In this role, you will define how the world measures scale-up fabric performance. You'll build the roofline models, benchmarks, and end-to-end workload studies that quantify our performance leadership, expose bottlenecks, and drive performance fine-tuning of real AI workloads on our fabric to inform product direction. Your data will directly shape architecture, firmware, and product decisions — and fuel the marketing narrative that positions Astera Labs at the center of AI connectivity. 

Key Responsibilities 

  • Performance Characterization & Benchmarking  
  • Establish theoretical and measured roofline models for Astera Labs' scale-up fabric across key performance metrics, defining the reference for all comparative testing. 
  • Build and maintain baseline performance benchmarks using industry-standard tools such as NVBandwidth and NCCL across a range of GPU configurations and switch topologies. 
  • Quantify the impact of differentiated Astera Labs AI fabric features (e.g., Hypercast, In-Network Computing) against baselines using both synthetic benchmarks and real inference workloads. 
  • Real Workload Analysis & Fabric Scalability  
  • Run end-to-end inference model workloads on target hardware to capture real-world performance beyond synthetic benchmarks, supporting architecture decisions and customer-facing demonstrations. 
  • Evaluate fabric performance as inference cluster size scales from 16 to 32 GPUs and beyond, identifying bottlenecks and building performance scaling models for state-of-the-art AI workloads. 
  • Design and execute head-to-head performance comparisons against competing fabric switch solutions to produce data-driven differentiation evidence. 
  • Test Infrastructure & Automation  
  • Design, build, and maintain automated lab infrastructure including test execution pipelines, traffic generation tooling, and data collection and reporting systems. 
  • Enable repeatable, high-quality, and scalable performance measurements across all hardware configurations, reducing manual effort and accelerating the test cycle. 
  • Share infrastructure and playbooks with the Product Applications team to accelerate customer application development and issue resolution. 
  • Cross-Functional Impact & Innovation  
  • Partner closely with ASIC architecture, firmware, software, Product Definition, Product Applications, and Product Marketing teams to communicate findings, influence design decisions, and resolve performance-impacting issues. 
  • Serve as a key technical resource in the early evaluation of new fabric architectures, interconnect technologies (UALink, PCIe Gen 6/Gen 7, Ethernet, UEC), and AI/ML communication paradigms. 
  • Provide performance data, analysis, and live benchmark support for key customer engagements and industry events; produce clear, audience-appropriate performance reports, technical briefs, and marketing collateral, and maintain living documentation in Confluence. 

Basic Qualifications 

  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field. We welcome both recent graduates with strong, directly relevant project, research, or internship experience and candidates with 2–5 years of industry experience in performance or systems engineering. 
  • Hands-on experience running AI/ML workloads on GPU clusters — including benchmarking, performance analysis, and fine-tuning of workloads across clusters of GPUs or accelerators. This can come from industry, research, or substantial academic projects. 
  • Demonstrated ability to debug and root-cause system-level performance issues across hardware, firmware, software, and network boundaries. 
  • Excellent fundamental knowledge of compute algorithms, parallel algorithms, and AI/ML algorithms and workloads. 
  • Strong working knowledge of computer systems, GPU systems, and datacenter networking — including PCIe and Ethernet fundamentals. 
  • Working knowledge of GPU and CPU software stacks (e.g., CUDA, MPI, collective communication libraries, drivers, and OS-level performance tooling). 
  • Proficiency in scripting and automation (e.g., Python) to build test pipelines and analyze large performance datasets. 

Preferred Qualifications 

  • MS or PhD in Computer Engineering, Computer Science, Electrical Engineering, or a related field. 
  • Experience with scale-up fabrics and next-generation interconnects such as UALink, PCIe Gen 6/Gen 7, Ethernet, or UEC. 
  • Deep understanding of modern inference and training workloads (LLMs, MoE, recommender systems) and their communication patterns. 
  • Experience developing roofline models and competitive performance analyses for switching, networking, or accelerator silicon. 
  • Excellent written and verbal communication skills, with the ability to translate deep technical findings into concise executive summaries and customer-facing narratives. 

Salary range is $135,000 to $170,000 depending on experience, level, and business need. This role may be eligible for discretionary bonus, incentives and benefits. 

We know that creativity and innovation happen more often when teams include diverse ideas, backgrounds, and experiences, and we actively encourage everyone with relevant experience to apply, including people of color, LGBTQ+ and non-binary people, veterans, parents, and individuals with disabilities.

Skills Required

  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or related technical field
  • Hands-on experience running AI/ML workloads on GPU clusters including benchmarking, performance analysis, and fine-tuning
  • Demonstrated ability to debug and root-cause system-level performance issues across hardware, firmware, software, and network boundaries
  • Excellent fundamental knowledge of compute algorithms, parallel algorithms, and AI/ML algorithms and workloads
  • Strong working knowledge of computer systems, GPU systems, and datacenter networking including PCIe and Ethernet fundamentals
  • Working knowledge of GPU and CPU software stacks (e.g., CUDA, MPI, collective communication libraries, drivers, and OS-level performance tooling)
  • Proficiency in scripting and automation (e.g., Python) to build test pipelines and analyze large performance datasets
  • MS or PhD in Computer Engineering, Computer Science, Electrical Engineering, or related field
  • Experience with scale-up fabrics and next-generation interconnects such as UALink, PCIe Gen 6/Gen 7, Ethernet, or UEC
  • Deep understanding of modern inference and training workloads (LLMs, MoE, recommender systems) and their communication patterns
  • Experience developing roofline models and competitive performance analyses for switching, networking, or accelerator silicon
  • Excellent written and verbal communication skills for executive summaries and customer-facing narratives
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, CA
148 Employees
Year Founded: 2017

What We Do

Astera Labs Inc., a fabless semiconductor company headquartered in the heart of California’s Silicon Valley, is a leader in purpose-built connectivity solutions for data-centric systems throughout the data center. Partnering with leading processor vendors, cloud service providers, seasoned investors, and world-class manufacturing companies, Astera Labs is helping customers remove performance bottlenecks in data-intensive systems that are limiting the true potential of applications such as artificial intelligence and machine learning. The company’s product portfolio includes system-aware semiconductor integrated circuits, boards, and services to enable robust CXL, PCIe, and Ethernet connectivity.

Similar Jobs

CoreWeave Logo CoreWeave

Senior Systems Engineer

Cloud • Information Technology • Machine Learning
In-Office
4 Locations
1450 Employees
182K-242K Annually

CrowdStrike Logo CrowdStrike

Senior Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
3 Locations
11000 Employees
160K-250K Annually

ServiceNow Logo ServiceNow

Staff Software Engineer

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Hybrid
Santa Clara, CA, USA
29000 Employees
191K-334K Annually

EROS Technologies Inc. Logo EROS Technologies Inc.

Performance Engineer

Artificial Intelligence • HR Tech • Information Technology • Consulting
In-Office
Sunnyvale, CA, USA
125 Employees

Similar Companies Hiring

Scrunch  Thumbnail
Artificial Intelligence • Information Technology • Marketing Tech • Software • SEO
Salt Lake City, Utah
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account