Founding Senior Applied AI Engineer (Agentic Systems)

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
Hybrid
185K-210K Annually
Senior level
Artificial Intelligence • Marketing Tech • Automation
The Role
Lead the design and implementation of production agentic AI systems for enterprise compliance. Responsibilities include multi-agent orchestration, RAG and memory systems, stateful Temporal workflows, evaluation and safety frameworks, OCR and document understanding pipelines, and AWS infrastructure. The role requires hands-on Python engineering, LLM application development, distributed systems expertise, and end-to-end ownership in an in-office San Francisco startup environment.
Summary Generated by Built In
Founding Senior Applied AI Engineer (Agentic Systems)

Our mission is to free humanity from meaningless work. We are building the System of Record for Enterprise Compliance, turning hours of manual, high-stakes marketing workflows into minutes of automated precision.

We are a small, high-density team of engineers and operators. Having proven our value with global brands, we are now at the inflection point where our technical architecture meets massive scale. This is a "rocket ship" moment: we are moving beyond simple automation into a world of vision-first, agentic workflows that solve the problems generic frontier models cannot.

The Role: Architect of the Agentic Brain

As Senior Applied AI Engineer, you’ll own the technical moat that makes us stand alone: the safety dataset and deterministic orchestration that prevents enterprise brands from ever trusting generic AI with compliance. You’re not building ‘better automation’—you’re building the category-defining infrastructure that gives us a 3-year lead in a market where second place doesn’t exist.

You will lead the technical evolution of our agent-driven architecture: defining the role of each agent, designing multi-step reasoning flows, integrating tools, memory, and retrieval systems, and optimizing for accuracy, determinism, and trust. Your work will directly determine whether enterprise customers can rely on Puntt for legal and brand compliance at scale.

This is a hands-on, in-office role in San Francisco, working closely with a small, senior team to build systems where correctness matters more than demos.

What You’ll Own

Agentic Reasoning & Orchestration

  • Design and evolve multi-agent LLM systems that decompose complex review tasks into reliable, auditable steps.

  • Define agent responsibilities, hand-offs, and termination conditions to minimize reasoning drift and maximize consistency.

Context, Retrieval & Memory Systems

  • Architect retrieval pipelines using RAG, structured memory, and emerging approaches like graph-based retrieval to provide agents with the right context at the right time.

  • Balance recall, precision, and latency across large knowledge bases (brand guidelines, regulations, historical decisions).

Stateful, Asynchronous Workflows

  • Own long-running, fault-tolerant workflows using Temporal (or similar), ensuring retries, versioning, and determinism across non-deterministic model calls.

  • Treat agent orchestration as a distributed systems problem: managing state, failures, and observability.

Evaluation, Safety & Reliability

  • Build evaluation frameworks that go beyond “it looks right,” using statistical metrics, gold labels, and automated regression testing to prove system reliability.

  • Prioritize correctness and trust, especially in high-risk legal and compliance scenarios.

Asset Understanding Pipeline

  • Collaborate on image and document preprocessing (OCR, layout analysis, VLMs) to ensure downstream agents receive structured, machine-readable context.

  • Focus on practical understanding, not computer vision research.

End-to-End Ownership

  • Move fluidly between Python-based LLM services, retrieval pipelines, and AWS infrastructure to ship reliable systems end-to-end.

Who You Are: The Hybrid Systems Builder

You are someone who enjoys building real systems with LLMs, not just experimenting with them.

  • Strong Engineering Foundation

You have 5–7+ years of experience building production systems and understand core CS concepts—data structures, concurrency, failure modes, and tradeoffs.

  • Experienced with LLM-Driven Systems

You’ve spent 1–2+ years building with large language models in real applications: tool use, function calling, structured outputs, and multi-step reasoning.

  • Agentic & Retrieval-First Thinker

You’re comfortable designing systems that combine LLMs with RAG, memory, graph-based context, and external tools rather than relying on a single prompt.

  • Systems-Oriented

You see multi-agent orchestration as a distributed systems challenge—latency, retries, observability, and consistency all matter.

  • Comfortable with Ambiguity

You thrive in an early-stage environment where problems are underspecified and the best solution doesn’t exist yet.

Technical Requirements

Must-have:

  • 5–7+ years of professional engineering experience, with a strong record of shipping production systems

  • 1–2+ years building with LLMs in real applications (not just experimentation)

  • Expert Python experience

  • Hands-on experience designing RAG systems, vector search, embeddings, and structured retrieval

*Preferred:

  • Experience with LLM orchestration frameworks (e.g., LangGraph, CrewAI, or custom orchestration layers)

  • Experience with stateful workflow orchestration (Temporal a plus)

  • Experience operating AI systems on AWS (Lambda, S3, Bedrock, SageMaker, etc.)

  • Strong TypeScript experience*

  • Bonus: experience with OCR, document parsing, or VLMs

Why Join Puntt

  • Small Team, Real Ownership

You will be a foundational technical leader shaping how the system works, not just implementing tickets.

  • High-Impact, High-Trust Domain

You’re building AI systems where correctness matters—and where most “generic AI” solutions fail.

  • Speed Without Chaos

We ship quickly, but we care deeply about system design, evaluation, and long-term reliability.

Skills Required

  • 5–7+ years of professional engineering experience building and shipping production systems
  • 1–2+ years building LLM-powered applications in production
  • Expert-level Python experience
  • Hands-on experience designing RAG systems, vector search, embeddings, and structured retrieval
  • Experience with LLM orchestration frameworks such as LangGraph, CrewAI, or custom orchestration layers
  • Experience with stateful workflow orchestration; Temporal experience is a plus
  • Experience operating AI systems on AWS, including Lambda, S3, Bedrock, or SageMaker
  • Strong TypeScript experience
  • Experience with OCR, document parsing, or vision-language models
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
11 Employees

Similar Jobs

InnoPeak Technology Logo InnoPeak Technology

Artificial Intelligence Engineer

Mobile • Other • Manufacturing
In-Office
Palo Alto, CA, USA
65 Employees
100K-200K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
197K-246K Annually

Edison Scientific Logo Edison Scientific

Artificial Intelligence Engineer

Artificial Intelligence • Software • Database
In-Office
San Francisco, CA, USA
47 Employees
160K-280K Annually

Factory Logo Factory

Software Engineer

Artificial Intelligence • Software
In-Office
San Francisco, CA, USA
125 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account