Senior Software Engineer, TorchPass Product Engineering

Posted Yesterday
Be an Early Applicant
Palo Alto, CA, USA
In-Office
150K-230K Annually
Senior level
Software
The Role
Build and own customer-facing installation, configuration, observability, documentation, and deployment capabilities for TorchPass. Debug production infrastructure, support customer evaluations, distinguish product defects from environment issues, improve tests and documentation based on field feedback, define cluster requirements and guarantees, and extend the product to new schedulers and job launchers. The role requires strong Python, Go, Kubernetes or Slurm, infrastructure operations, customer support, and technical documentation skills.
Summary Generated by Built In
About Clockwork Systems

Clockwork.io – Software Driven Fabrics to increase GPU cluster utilization

Clockwork Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the foundations of distributed computing. As AI workloads grow increasingly complex, traditional infrastructure struggles to meet the demands of performance, reliability, and precise coordination. Clockwork is pioneering a software-driven approach to AI fabrics by delivering cross-stack observability to catch and quickly resolve problems, workload fault tolerance to keep jobs running through failures, and performance acceleration that dynamically routes and paces traffic to avoid congestion.

To learn more, visit www.clockwork.io.
About the Role

TorchPass keeps distributed PyTorch training running through hardware and network failures. The Product Engineering team makes it something customers can install, trust, and run on their own clusters.

We are looking for an experienced engineer who has turned working infrastructure software into a product that other people can deploy and operate. You will own the parts of TorchPass that customers touch, and you will be one of the engineers they hear from when something does not work.

This is a product and deployment role. It does not require GPU runtime, NCCL, or kernel experience.

What You Will Do
  • Build and own the customer-facing surfaces of TorchPass: installation, configuration, and the views that show customers what the product is doing
  • Keep product documentation accurate with every release
  • Work directly with customers and our solutions engineers during evaluations: read logs, separate product defects from environment issues, and say clearly what happened
  • Turn what we learn in the field into fixes, tests, and documentation
  • Define, together with the core engineering team, what TorchPass needs from a customer's cluster and what it guarantees in return
  • Extend TorchPass to new schedulers and job launchers as customers need them
What We're Looking For
  • 4+ years building or operating infrastructure software in production
  • You have taken a system that worked for its authors and made it installable and supportable by people who did not build it
  • Strong debugging skills on production systems you did not write
  • Deep experience with Kubernetes or Slurm
  • Strong Python, and the ability to read and change Go
  • You use AI coding tools well, and you understand and stand behind the code you ship
  • You write documentation that people follow successfully
  • You are comfortable working with customers on technical problems
  • A degree in computer science, electrical engineering, or a related field, or equivalent experience
Preferred
  • Experience with GPU clusters or machine learning training infrastructure
  • Experience with release engineering: build pipelines, packaging, artifact delivery
  • Experience bringing up clusters across more than one cloud provider

Enjoy

  • Challenging projects.
  • A friendly and inclusive workplace culture.
  • Competitive compensation.
  • A great benefits package.
  • Catered lunch.

Compensation for this position will vary based on the skills and experience you bring, as well as internal equity considerations. For candidates hired at the posted level, the expected base salary range is $150,000 - $230,000. In addition to cash compensation, this role is eligible to participate in the company’s equity program, which may include stock options granted in accordance with the company’s equity plan and subject to approval and applicable vesting schedules.

Clockwork Systems is an equal opportunity employer. We are committed to building world-class teams by welcoming bright, passionate individuals from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, ancestry, religion, age, sex, sexual orientation, gender identity or expression, national origin, disability, or protected veteran status. We believe diversity drives innovation, and we grow stronger together.

Skills Required

  • 4+ years building or operating infrastructure software in production
  • Experience making infrastructure software installable and supportable by users who did not build it
  • Strong debugging skills on production systems not previously authored
  • Deep experience with Kubernetes or Slurm
  • Strong Python skills
  • Ability to read and modify Go
  • Ability to use AI coding tools effectively and take responsibility for shipped code
  • Ability to write accurate, usable technical documentation
  • Comfort working directly with customers on technical problems
  • Degree in computer science, electrical engineering, or a related field, or equivalent experience
  • Experience with GPU clusters or machine learning training infrastructure
  • Release engineering experience, including build pipelines, packaging, or artifact delivery
  • Experience bringing up clusters across multiple cloud providers
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
26 Employees
Year Founded: 2018

What We Do

Founded in 2018 by a team from Stanford University, Clockwork's technology enables time-sensitive applications in areas such as financial trading, high-tech, and online gaming. Being software-based, its solutions can run anywhere: in on-premises data centers, public clouds, or hybrid environments. Taking aim at the 'clockless architecture' prevalent in distributed systems and networks, Clockwork.io aims to redefine a large part of the way these technologies (which underlie the cloud) are currently practiced. We're hiring software engineers, deployment engineers and sales. Send email to [email protected]

Similar Jobs

Micron Technology Logo Micron Technology

Director, CMBU Product Management

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
2 Locations
45000 Employees
170K-353K Annually

Micron Technology Logo Micron Technology

Application Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
San Jose, CA, USA
45000 Employees
204K-347K Annually

Micron Technology Logo Micron Technology

Program Coordinator

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
San Jose, CA, USA
45000 Employees

Circle Logo Circle

Senior Site Reliability Engineer

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
In-Office or Remote
San Francisco, CA, USA
1050 Employees
153K-205K Annually

Similar Companies Hiring

Ford Energy Thumbnail
Automotive • Software • Energy • Utilities • Manufacturing • Renewable Energy
US
55 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
70 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account