Senior Founding AI Software Engineer

Reposted Yesterday
Hiring Remotely in San Francisco, CA, United States
In-Office or Remote
190K-270K Annually
Senior level
Artificial Intelligence • Software • Biotech • Pharmaceutical
Building the future of biotech by anticipating FDA approval probability.
The Role
Build and operate production AI systems for FDA regulatory intelligence, including evaluation pipelines, document ingestion and retrieval, multi-agent workflows, data pipelines, cloud infrastructure, and secure multi-tenant services. Collaborate with former FDA experts to translate domain judgment into product requirements and evaluation cases. Own architecture, reliability, observability, cost, security, and rapid 0-to-1 delivery across the stack.
Summary Generated by Built In

About the role

This is a rare and exciting opportunity for a hands-on builder who wants to take ownership and build 0 to 1 at a VC-backed, fast-growing startup. Build with former FDA decision-makers to define the future of biotech industry standards.

You don't have to know FDA or biotech, but you believe in vertical AI solutions and want to push the boundary of what is possible. Work across our entire stack from document ingestion to agents and evals, build 0 to 1 and grow with a young team. We're looking for a Senior Founding AI Software Engineer who is passionate about solving challenging problems in a highly regulated industry.

Headquartered in San Francisco, Deffai is an AI-native VC-backed seed-stage startup building an AI platform for simulating FDA regulatory intelligence, built together with former FDA reviewers. We help therapeutics companies and their investors anticipate regulatory outcomes.

Backed by top-tier tech investors in the SF Bay Area, with advisors from GitHub, OpenAI, and Mercor.

You get market standard comp, equity, viable path to Head of Engineering or beyond.

You will work closely with 2 engineers on the team, build, ship quickly and shape our product vision.

If you have a strong interest, qualify for some but not all of what’s mentioned, please reach out with a short note! We manually review every application and we might make an exception.

Key Responsibilities

  1. Evals & Benchmarks: Build the eval pipeline for model benchmarking and regression suites. Use them to gate prompt, model and agent changes, measure precision and recall against expert judgment, and benchmark our performance with frontier models.
  2. Document Ingestion & Retrieval: Build ingestion and retrieval over diverse, long, and complex regulatory documents (scanned PDFs, Word files, tables, slide decks): parsing, structure-aware chunking, hybrid search, reranking, context assembly under token budgets, and knowledge graphs. Keep every answer traceable to a cited source, with provenance and versioning.
  3. Agents: Build and operate our multi-agent review workflows as they move to a shared LangGraph engine: tool use, memory, tracing, and handling cost, latency, and failures in long-running runs.
  4. Expert Collaboration: Work with former FDA experts to accelerate their regulatory review workflows. Turn their feedback into product changes and new evaluation cases.
  5. Build & Ship: Own architectural decisions that balance innovation, scalability, cost, and long-term maintainability in a high-stakes domain. Review AI-generated code for what a diff does not show: migration locks, environment boundaries, and tenant scope.
  6. Data Engineering: Work with the team on the data pipelines behind our product: document ingestion, background jobs, and the PostgreSQL schemas. Make every pipeline idempotent and observable.
  7. Infrastructure: Support the team on our AWS infrastructure: Terraform-managed ECS services, CI/CD, and separate development and production environments.
  8. Security & Tenant Isolation: Support the team with tenant isolation, role-based access control, encryption, audit logging, and data retention controls in every service that handles customer documents.

Who you are

  1. Cracked: You love to build.
  2. Strong opinions: You have strong opinions and strong product taste, and you're not afraid of pushing back.
  3. Relentless: You are tough, can handle the speed and execute quickly.
  4. Execution-Oriented: You move fast and focus on solving real problems.
  5. Clear Communicator: You collaborate well across functions and push decisions forward.

Requirements

  1. 5+ years building and operating production software, including owning the architecture of a production system
  2. Experience building evaluations against expert judgment, such as golden datasets and precision and recall, and using them to decide what ships
  3. Experience shipping LLM systems to production, such as retrieval over complex documents or agent workflows
  4. Daily use of AI coding tools, with the judgment to catch what they get wrong in migrations, infrastructure changes, and data-access code
  5. Strong Python and data engineering experience: schema design, zero-downtime migrations, and reliable pipelines and background jobs on a relational database
  6. Experience running production systems on a major cloud with infrastructure as code, CI/CD, and separate development and production environments

Nice to have

  1. Worked directly with domain experts (clinicians, lawyers, scientists) to turn their judgment into requirements and test cases
  2. Retrieval techniques: hybrid search, reranking, parsing scanned PDFs and tables, or knowledge graphs (GraphRAG or similar)
  3. Agent frameworks and runtimes: LangGraph, MCP, Claude Agent SDK, or sandboxed agent execution
  4. Post-training or fine-tuning (SFT, preference tuning, LoRA), and judgment on when it beats prompt and retrieval work
  5. Experience building secure multi-tenant systems: tenant isolation, access control, and audit logging
  6. AWS (ECS, RDS, S3) and Terraform

Our stack

You don't need prior experience with every tool in our stack. We value strong engineering fundamentals and the ability to take ownership of production systems.

  • Frontend: React, TypeScript, Vite, Tailwind CSS, Vercel AI SDK
  • Backend and persistence: Python, FastAPI, SQLAlchemy, PostgreSQL, Alembic
  • Cloud: AWS ECS/Fargate, ECR, RDS, S3, Application Load Balancer, Secrets Manager; Cloudflare for DNS
  • Delivery and infrastructure: Docker, GitHub Actions, Terraform and OpenTofu
  • AI and orchestration: Claude via the Anthropic API; application-managed skill packs, tool use, retrieval, streaming, and evaluation workflows
  • Data and documents: FDA corpus ingestion and SQL analytics, PDF parsing with PyMuPDF, DOCX processing and report generation with python-docx

Interview process

  1. 30min intro with CEO & cofounder
  2. 40min technical interview with one of our engineers, including a short hands-on exercise
  3. Short paid take-home (5-8 hr), followed by a walkthrough where you explain your thought process
  4. Final round: meet one of our former FDA decision makers

We aim to decide within 2 weeks of your final round. Don't miss this rare opportunity to define the next generation of medicine approvals.

We currently do not provide visa sponsorship.

Skills Required

  • 5+ years building and operating production software, including owning production system architecture
  • Experience building evaluations against expert judgment, including golden datasets, precision and recall, and release gating
  • Experience shipping LLM systems to production, such as retrieval over complex documents or agent workflows
  • Daily use of AI coding tools and judgment to identify errors in migrations, infrastructure changes, and data-access code
  • Strong Python and data engineering experience, including schema design, zero-downtime migrations, reliable pipelines, background jobs, and relational databases
  • Experience operating production systems on a major cloud with infrastructure as code, CI/CD, and separate development and production environments
  • Experience working directly with domain experts such as clinicians, lawyers, or scientists to turn judgment into requirements and test cases
  • Experience with hybrid search, reranking, scanned PDF and table parsing, or knowledge graphs
  • Experience with LangGraph, MCP, Claude Agent SDK, or sandboxed agent execution
  • Experience with post-training or fine-tuning, including SFT, preference tuning, or LoRA
  • Experience building secure multi-tenant systems with tenant isolation, access control, and audit logging
  • Experience with AWS services including ECS, RDS, and S3, and Terraform
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
8 Employees
Year Founded: 2025

What We Do

Headquartered in San Francisco, Deffai is an AI-native VC-backed seed-stage startup building an AI platform for simulating FDA regulatory intelligence, built together with former FDA reviewers. We help therapeutics companies and their investors anticipate regulatory outcomes. We are ambitious builders with diverse technical backgrounds building vertical AI solutions in a deeply regulated space. Work closely with former FDA decision-makers to redefine the future of biotech approvals. Backed by top-tier tech investors in the SF Bay Area, with advisors from GitHub, OpenAI, and Mercor. Are you an ambitious builder who wants to do more and push the boundary of what's possible? Join our team! Size of the opportunity: FDA accounts for 20% of the U.S. economy alone. Food, drug, and cosmetics. Beyond the U.S., other countries and regions including Canada, Europe, and Japan all have their own FDA with similar frameworks.

Why Work With Us

We are ambitious, highly technical, and collaborative. We love it when you have strong opinions. You will get to work across our entire stack and directly with CEO. This is a rare and exciting opportunity for a hands-on builder who wants to take ownership and build 0 to 1 at a VC-backed, fast-growing startup.

Gallery

Gallery

Similar Jobs

Remote
USA
142 Employees
146K-250K Annually

Shield AI Logo Shield AI

Senior Principal Autonomy Engineer (R6162)

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software • Defense Technology
Remote
USA
330K-500K Annually
In-Office or Remote
2 Locations
175633 Employees
133K-284K Annually

General Motors Logo General Motors

Account Manager

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
United States
165000 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account