About the company
Our client is a fast-growing technology company.
The role
Raydar is recruiting for this role on behalf of our client. Take ownership of the platform that lets research teams train and run large models efficiently. You will bridge researchers and compute, shaping distributed systems, raising hardware efficiency and delivering the data flows that power experiments. This is a senior, high-autonomy position.
What you'll do
- Architect and grow the systems that spread model training over many accelerators, with attention to memory efficiency.
- Create easy-to-use internal tools and workload scheduling so researchers can iterate quickly and recover gracefully from failures.
- Engineer fast data ingestion and transformation pipelines that handle very large volumes of varied sensor and media data.
- Prepare trained models for fast, responsive deployment in live settings using compression and compilation methods.
- Investigate performance in depth, tracing slowdowns in compute, storage and memory to get more from a growing hardware fleet.
- Work side by side with researchers to shorten the time from experiment to production.
Requirements
What we're looking for
- 7+ years in ML or data infrastructure, including technical leadership on HPC or ML infrastructure projects.
- A track record of running production ML or data platforms within a highly skilled engineering team.
- Deep hands-on knowledge of PyTorch.
- Strong command of accelerator performance tuning, serving efficiency and observability.
- Practical experience with distributed training libraries.
- Prior experience joining a young company as one of its first infrastructure engineers.
- Experience building systems for multimodal models such as video, audio or other media.
- Genuine enthusiasm for robotics.
- Willingness to work on-site five days a week.
Bonus points
- Robotics experience in a startup or enterprise setting.
Benefits
Compensation and benefits
- Base salary: USD 220,000 to 350,000 per year
- Equity
- Flexible time off
- Relocation support
Location and work model
- Redwood City, CA, United States
- On-site, 5 days per week in office
- Full-time
Skills Required
- 7+ years of experience in machine learning or data infrastructure
- Technical leadership on HPC or machine learning infrastructure projects
- Experience running production machine learning or data platforms within a highly skilled engineering team
- Deep hands-on knowledge of PyTorch
- Strong knowledge of accelerator performance tuning, serving efficiency, and observability
- Practical experience with distributed training libraries
- Experience joining a young company as one of its first infrastructure engineers
- Experience building systems for multimodal models, such as video, audio, or other media
- Genuine enthusiasm for robotics
- Willingness to work on-site five days per week
- Robotics experience in a startup or enterprise setting
What We Do
Raydar is a talent acquisition and business consulting firm that connects world-class and emerging talent with growing organizations. It supports companies through team development, strategic hiring, and customized growth solutions, helping clients recruit roles such as engineers, product managers, executives, legal counsel, and quantitative traders. Raydar focuses on understanding each organization’s needs, culture, and long-term goals to build high-impact teams.








