About the company
Our client is an AI technology company.
The role
Raydar is recruiting for this role on behalf of our client. Own the measurement of model quality from start to finish, defining metrics, building the tooling to compute them and turning the results into concrete improvements. The work centers on evaluating complex multi-component machine learning systems where no ground truth exists. You will carry projects from the initial question through to a written conclusion without handoffs.
What you'll do
- Define evaluation metrics and what improvement means for outputs that evolve over time, keeping those definitions current as methods change.
- Build data pipelines and test harnesses covering data ingestion, labeling, versioning, reruns and automated judging.
- Analyze run traces and results to find real failure modes rather than only what is easy to measure, and propose fixes.
- Develop simulation tooling that lets many scenarios be modeled and measured in parallel to speed up iteration.
- Own each evaluation loop end to end, from the question to the pipeline, rerun and final writeup.
Requirements
What we're looking for
- Experience designing, building and running evaluations for large language models from start to finish.
- Prior work in areas such as memory, personalization, agent evaluation or agent optimization.
- Published research as a first or co-author at a top machine learning venue, or open-source work of comparable quality.
- A university-level background in a quantitative discipline such as computer science or statistics.
- Production-quality Python, including testing, maintenance and regular releases.
- Working knowledge of PyTorch and common machine learning tooling and experiment tracking.
Bonus points
- Experience with data pipeline tooling such as Kafka and SQL.
Benefits
Compensation and benefits
- Base salary: USD 220,000 to 300,000 per year
- Equity
- Health, dental and vision coverage
- Retirement plan
- Paid time off
- Relocation support
Location and work model
- New York, NY, United States
- On-site, 5 days per week in office
- Full-time
Skills Required
- Experience designing, building, and running evaluations for large language models end to end
- Prior experience with memory, personalization, agent evaluation, or agent optimization
- First- or co-author published research at a top machine learning venue, or comparable-quality open-source work
- University-level background in a quantitative discipline such as computer science or statistics
- Production-quality Python experience, including testing, maintenance, and regular releases
- Working knowledge of PyTorch and common machine learning tooling and experiment tracking
- Experience with data pipeline tooling such as Kafka and SQL
What We Do
Raydar is a talent acquisition and business consulting firm that connects world-class and emerging talent with growing organizations. It supports companies through team development, strategic hiring, and customized growth solutions, helping clients recruit roles such as engineers, product managers, executives, legal counsel, and quantitative traders. Raydar focuses on understanding each organization’s needs, culture, and long-term goals to build high-impact teams.









