RESEARCHER, AGENTS FOR AUTOMATED DISCOVERY

Posted 2 Days Ago
Be an Early Applicant
San Francisco, CA, USA
In-Office
Senior level
Artificial Intelligence • Information Technology • Machine Learning • Software
The Role
Research and develop autonomous multi-agent systems that plan, decompose tasks, choose experiments, evaluate outputs, and recover from failures. Design methods, coordination patterns, and evaluation frameworks for open-ended research workflows; run rigorous experiments, ship methods to production, and drive team research direction.
Summary Generated by Built In

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You'll be researching the agents at the core of our work: multi-agent systems that conduct automated machine learning research and discovery. You'll design how these agents plan, decompose problems, choose what to try next, evaluate their own outputs, and recover from mistakes.

This is a deeply open-ended research role. The benchmarks for agents that do real research don't exist yet, and inventing them is part of the job. You'll move between method design, careful experimentation, building evaluation frameworks, and shipping into production. Real autonomy, real ownership, and the corresponding responsibility for choosing well.

WHAT YOU'LL DO

  • Design methods that improve how our agents plan, decompose tasks, use tools, manage context, and recover from failures across long-horizon research workflows

  • Develop multi-agent coordination patterns: how multiple agents share context, divide labor, supervise each other, and combine their outputs

  • Build and maintain evaluation frameworks for agent capability on open-ended tasks (the kind where the right answer isn't pre-specified)

  • Run rigorous experiments to characterize what works, what doesn't, and why: controls, ablations, statistical significance

  • Co-design agent architectures with engineering teammates; ship the most promising methods into production

  • Read deeply across the agentic ML, planning, RL, and tool-use literature; bring useful work from outside in

  • Share findings internally so the rest of the team builds on them

  • Help shape research direction across the team: agentic research taste compounds when discussed openly

WHAT WE'RE LOOKING FOR

  • Strong track record of ML research with focus on agents, RL, LLMs, planning, tool use, or multi-agent systems

  • 5+ years of hands-on research experience in industry or academia

  • Comfort designing experiments and running them end-to-end at scale

  • Track record of building evaluation frameworks for capabilities that aren't easily benchmarked

  • Bias toward shipping research, not handing it off

  • Strong written communication: you can compress a result into a paragraph that changes what someone else does next

  • Comfort with ambiguity: open-ended problems without fixed benchmarks are the work, not a frustration

  • Published research at NeurIPS, ICML, ICLR, COLM, RLC, or comparable venues

NICE TO HAVE

  • PhD in ML, statistics, CS, or adjacent

  • Published research on agentic systems, tool use, long-horizon planning,

  • multi-agent coordination, or self-improvement methods

  • Open-source contributions in the agentic ML ecosystem (coding agents,

  • research assistants, autonomous workflows)

  • Experience with reasoning models, chain-of-thought / scratchpad methods,

  • or supervised fine-tuning for agentic behaviors

  • Background in evaluation methodology for capabilities that don't have

  • established benchmarks

THIS ROLE IS PROBABLY NOT FOR YOU IF

  • You want to focus on a single stable benchmark: our agents work on open-ended problems and the targets shift

  • You prefer to keep research paper-only; these agents need to actually work- You'd rather work alone than share research taste openly with a small team

Skills Required

  • Strong track record of ML research focused on agents, RL, LLMs, planning, tool use, or multi-agent systems
  • 5+ years of hands-on research experience in industry or academia
  • Comfort designing experiments and running them end-to-end at scale
  • Track record of building evaluation frameworks for capabilities that aren't easily benchmarked
  • Bias toward shipping research and integrating methods into production
  • Strong written communication ability to concisely summarize results
  • Comfort with ambiguity and working on open-ended problems without fixed benchmarks
  • Published research at NeurIPS, ICML, ICLR, COLM, RLC, or comparable venues
  • On-site work in San Francisco
  • PhD in ML, statistics, CS, or adjacent
  • Published research on agentic systems, tool use, long-horizon planning, multi-agent coordination, or self-improvement methods
  • Open-source contributions in the agentic ML ecosystem (agents, research assistants, autonomous workflows)
  • Experience with reasoning models, chain-of-thought / scratchpad methods, or supervised fine-tuning for agentic behaviors
  • Background in evaluation methodology for capabilities that lack established benchmarks
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
8 Employees

What We Do

MakerMaker.AI is an innovative company building autonomous research agents focused on recursive self-improvement. They develop sophisticated multi-agent systems that can independently propose, run, and analyze complex machine learning experiments. By creating agents that have the ability to build other agents, MakerMaker.AI seeks to accelerate the development of artificial intelligence and push the boundaries of autonomous research in machine learning.

Similar Jobs

Notion Logo Notion

Talent Management

Artificial Intelligence • Productivity • Software
Hybrid
San Francisco, CA, USA
1000 Employees
230K-270K Annually

CrowdStrike Logo CrowdStrike

Threat Hunter III (Remote)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
100K-155K Annually

CrowdStrike Logo CrowdStrike

Sr. Intelligence Analyst (Remote)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
85K-120K Annually

Applied Systems Logo Applied Systems

VP, Agency Adoption & Engagement

Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Remote or Hybrid
United States
3079 Employees
164K-221K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account