Head of Research

Posted 4 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
50K-500K Annually
Senior level
Artificial Intelligence • Software • Automation
The Role
Lead and build an applied research org focused on agent planning, tool orchestration, multi-agent coordination, evaluation, and safety. Set research roadmap, run experiments and A/Bs, own evaluation stack and datasets, develop novel methods and reference implementations, and partner with engineering to productionize improvements that increase reliability and performance. Role is on-site in San Francisco.
Summary Generated by Built In
About Datawizz

Datawizz is building the agent workforce. We're a seed-stage team backed by Human Capital (SpaceX, Snowflake, Anduril) building a platform that helps enterprises build and deploy agents — with the right permissions, guardrails, and activity logging to actually ship within companies. Early customers are already automating several hours of their workday.

We're hiring people who want to build, not manage. If you want to work on hard problems with a small team that ships, this is the place.

The Role

As Head of Research, you’ll own Datawizz’s research agenda and build an applied research organization focused on making agents radically more reliable and capable. You’ll define the roadmap across agent planning, tool orchestration, evaluation, and safety; lead hands-on experimentation; and partner closely with engineering to productionize breakthroughs that drive measurable quality and reliability wins for customers.

You will:
  • Set the research strategy and roadmap for agent planning, tool orchestration, multi-agent coordination, and evaluation.

  • Build, lead, and mentor a high-caliber applied research team.

  • Design and run rigorous experiments (ablations, offline/online A/Bs), defining clear metrics for reliability, accuracy, and performance.

  • Own our evaluation stack: datasets, benchmarks, human-in-the-loop reviews, and reliability/safety assessments.

  • Develop novel methods (e.g., agent planning algorithms, tool selection strategies, multi-agent coordination, safety/alignment techniques) and ship reference implementations.

  • Collaborate with engineering to transfer research into production and measure real-world impact.

  • This role is in-office, 5 days/week, based in San Francisco.

You might be a great fit if you have experience with:
  • Leading applied ML/NLP research teams and shipping work into production at a startup or high-growth company.

  • LLM and agent internals: prompting strategies, tool use, planning/reasoning architectures, multi-agent systems, and evaluation methodology.

  • Building evaluation frameworks (task suites, synthetic data, human eval pipelines) and tying metrics to product outcomes.

  • Large-scale agent systems: orchestration frameworks, tool APIs, distributed execution, observability, and logging infrastructure.

  • Strong coding skills in Python and a bias toward hands-on experimentation and rapid iteration.

  • Data curation and labeling workflows, with attention to privacy, safety, and robustness.

  • Communicating research clearly and partnering cross-functionally with engineering and product.

  • (Nice to have) Publications or notable open-source contributions; patents; early-stage 0→1 experience.

Benefits
  • Competitive salary, based on experience level (Annual compensation range: $50,000-$500,000)

  • Meaningful equity

  • Opportunity to be a founding member of a growing company

Skills Required

  • Set research strategy and roadmap for agent planning, tool orchestration, multi-agent coordination, and evaluation.
  • Build, lead, and mentor an applied research team.
  • Design and run rigorous experiments (ablations, offline/online A/B tests) with clear reliability and performance metrics.
  • Own evaluation stack: datasets, benchmarks, human-in-the-loop reviews, and reliability/safety assessments.
  • Develop novel methods (agent planning algorithms, tool selection strategies, multi-agent coordination, safety/alignment techniques) and ship reference implementations.
  • Collaborate with engineering to transfer research into production and measure real-world impact.
  • Experience leading applied ML/NLP research teams and shipping work into production at a startup or high-growth company.
  • Deep knowledge of LLM and agent internals: prompting, tool use, planning/reasoning architectures, multi-agent systems, and evaluation methodology.
  • Experience building evaluation frameworks, human eval pipelines, and tying metrics to product outcomes.
  • Experience with large-scale agent systems: orchestration frameworks, tool APIs, distributed execution, observability, and logging infrastructure.
  • Strong coding skills in Python and bias toward hands-on experimentation and rapid iteration.
  • Experience with data curation and labeling workflows, with attention to privacy, safety, and robustness.
  • On-site work requirement: in-office 5 days/week in San Francisco.
  • Publications, notable open-source contributions, patents, or early-stage 0->1 experience.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
9 Employees
Year Founded: 2016

What We Do

Gamut is an artificial-intelligence software platform for building and running personal and enterprise AI agents in secure environments. Its agents operate autonomously in browsers and connected tools, handling research, scheduling, repetitive preparation, and other complete workflows rather than isolated tasks. The company provides controls such as permissions, guardrails, activity logging, and integrations so teams can deploy agents productively within their organizations.

Similar Jobs

Valency International Logo Valency International

Head of Research

Food • Agriculture • Chemical • Industrial
Hybrid
Berkeley, CA, USA
2600 Employees

Rwazi Logo Rwazi

Head of Research & Development (R&D)

Big Data • Information Technology • Analytics • Business Intelligence
In-Office or Remote
7 Locations

Collinear AI Logo Collinear AI

Head of Research

Artificial Intelligence • Machine Learning • Software • Generative AI
In-Office
Sunnyvale, CA, USA
250K-400K Annually

Traverse Logo Traverse

Head of Research

Artificial Intelligence • Big Data • Machine Learning • Software
In-Office
San Francisco, CA, USA

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account