AI Engineer

Posted 10 Hours Ago
Metropolitan, CA, USA
In-Office
250K-450K Annually
Expert/Leader
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Product.ai is the truth layer for commerce.
The Role
Own and evolve the production agent harness for the CEO office: design long-lived agent runs, verification systems, liveness checks, and token/model-routing. Build regression corpora, oracle-separated checkers, CI gates, and deterministic companions so automations never fail silently. Work directly with the founder, ship hardened systems, and measure falsifiable outcomes for a high-agency, production-critical AI platform.
Summary Generated by Built In
A builder with a high technical bar whose leverage is judgment, not keystrokes - in the Office of the CEO.

Product.ai is the verified truth layer for shopping - the intelligence that tells you what's actually true about a product, including when not to buy. Our first proof at scale is SimplyCodes, the code verification service: roughly $22M in revenue at ~60% margins. Profitable. Bootstrapped. No outside investors. No board. A small team of fewer than twenty operators, outbuilding companies 10× our size.

Strong people find us and keep finding us - they apply over months and years, because the field moves fast and the exact profile we need moves with it.

Why This Role Exists

Most of our software is now written by AI agents. So the job that matters is no longer typing the code - it's deciding what to build, designing the systems the agents run inside, and knowing how you'll prove the work is correct. You'll be a builder closer to a product engineer than a heads-down coder: your leverage is judgment and taste, not typing speed.

Your first surface is the agent harness behind the Office of the CEO - the live automation that lets fewer than twenty operators move like a company many times that size. A recruiting-evaluation pipeline scoring 30+ candidates a day across open roles. A merchant-discovery pipeline landing ~1,200 merchants a day. The content, data, and ops automation underneath it all. This year these crossed a threshold: agent runs that go 1-4 hours unattended became a normal unit of work for us, and that fleet now deserves a dedicated owner.

You won't stay boxed into one surface. A lot of the work is finishing what's nearly done: the founder starts a system, it runs but isn't yet hardened or owned, and you take it the last mile and own it from then on. You work inside his build loop, not beside it - the harness is where you start, not the ceiling of what you'll touch. This is a generalist builder's seat, working directly with the founder.

The System You'll Need to Model

  • A fleet of production automations whose failure mode is silent death. Pipelines here rarely fail loudly - they stop, and the cost accrues invisibly until someone notices days later. The real engineering problem is liveness: designing alarms and deterministic checks so that no automation in the fleet can die unnoticed.
  • Long-lived agent runs that go 1-4 hours unattended. They hold together not because someone watches them, but because they run on architectural law (the rules an agent run is bound by), a fuel budget, and verification built in. You design what governs a run, what it's allowed to spend, and what proves it worked.
  • Verification the agent cannot author. A generative model cannot reliably grade its own output, so a verifier that shares the generator's context will launder its own mistakes. The architecture is external truth anchors, regression corpora, and oracle-separated checkers - a separate judge that never sees what the builder saw. This separation is the whole game, and it's also the company's thesis: verified truth a model can't fake.
  • Token spend judged by what it moved, not what it cost. Every run is instrumented for the outcome it produced, and budget gets redirected toward what's working while the run is still going. We're quality-maximalist: the expensive thing is a redo cycle, never tokens.
  • Cortex - the shared AI brain you work inside. It isn't a tool you reach for now and then; it's the operating system that runs the whole company, and it's the same intelligence family we sell to the world. Every operator here works through governed AI sessions on it, and it answers its own questions from a base of 8,600+ documents. The automations you build stand on it, feed it, and are bound by the same architectural law your own work is. Nobody else runs their company on the product they sell.
  • Architecture that moves weekly. We built this harness before unattended runs were even possible at scale, and the ground keeps shifting. You'll model where it's going and act without waiting for a brief.


If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.

What You Will Own

  • The agent harness for the Office of the CEO. Recruiting evaluations, merchant discovery, ops automation - the run designs, the runtime they execute in, and the architectural law that governs them. When a new automation is needed, you decide how it runs, what governs it, and what proves it correct.
  • Verification that scales as the work compounds. Ad-hoc human review collapses somewhere around 100-150 artifacts a day, and we're heading straight through that ceiling. You build what replaces it: regression corpora, oracle-separated checkers, sampling protocols, and escalation paths that put a human in the loop only where real judgment is needed.
  • A liveness check on every automation. Each system ships with a deterministic companion that proves it's alive and correct - you own the standard and the coverage. Nothing runs in production without one; nothing dies unnoticed.
  • Token economics and model-routing for the fleet. You decide which model runs which job and why, pointing real compute at business outcomes and adjusting the spend mid-run. You'll own falsifiable outcomes - each with an evidence test a stranger could run.
  • Your seat, on a co-signed charter. This is the model we run here: within your first quarter, you and the founder co-sign a seat charter - one machine-checkable number that proves the seat is working, and a written split of what you decide on your own versus what you bring to him first. You own that number, and the authority that comes with it.


Who You Are

You form working models of running systems on your own. You can read a pipeline you've never seen and sketch its failure modes the same day, notice where your model is wrong, and update fast. When a build comes out wrong, you fix the spec that produced it, not just the symptom in front of you. You write clearly, because clear writing is evidence of clear thought - and here, what you write becomes the law your agents execute.

You move between architecture and implementation without getting stuck at either altitude: designing a verification gate in the morning, shipping it alarmed and instrumented by the afternoon. You have strong, earned opinions about guardrails, evaluations, and agent behavior, and you make good calls in the gray instead of queuing questions. High agency is your resting state.

You can do this job by hand, and prove it - that depth of craft is exactly what lets you direct agents and still trust the result. You've built automations that ran in production for months, and you can say precisely how you knew when one was failing. Evaluation harnesses, CI gates that actually block, data pipelines with verification companions, agent systems with regression suites. You treat agents as leverage you verify, not autocomplete you trust - depth is what lets you trust the verdict. We care about the artifact and the reasoning behind it far more than where you built it or what's on your diploma.

Who this isn't for. This isn't the seat for everyone. It's wrong if your code is whatever the model handed you and you couldn't say why it's right - directing agents without depth of your own breaks down fast here. It's wrong if you're comfortable letting an agent grade its own work. It's wrong if you work best as a watched assistant or want a ticket queue to execute and report done. It's wrong if you think mainly in projects, timelines, and programs - the steering here happens in real time, mid-run. And it's wrong if you mostly optimize for logos on a resume. You'll be happiest if you ship the verification companion with the feature, fix the spec instead of the symptom, and want your visibility to come from registered architecture decisions and outcomes moved - not hours, meetings, or activity.

How We Evaluate

We don't run traditional systems-engineering interviews.

  • Video screen. Brief and async: 5-6 questions, about 15 minutes total, done whenever works for you.
  • Calls with company stakeholders. Short conversations with key members of the team.
  • Conversation with the founder. Chemistry and comprehension - can you model the system you just read about?
  • Paid work trial. A paid one-to-two-week engagement - real work in our real environment. We watch how you ground yourself, whether you write the spec before the build, how you verify what your agents produce, and whether your self-assessment is honest. It is paid because it is real work, and because that respects your time.


  • If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the artifact and the reasoning, not the pedigree.

    Compensation & Ownership

    Total first-year comp: $375,000-$450,000 (base + equity + profit sharing). Base: $250,000-$300,000. We set the number from what you've actually built, and every hire ships from day one.

    Profits Interest Units (PIUs) - Class B Membership Interests at $0 strike, real ownership from day one, capital-gains treatment; annual pro-rata profit sharing from free cash flow; annual tender liquidity; 100% family premium coverage; and an effectively unlimited token budget, steered by ROI, never capped.

    This is a partnership, not a salary line. The model is built to mint partners - when the company wins, you win, in real, liquid dollars, every year.

    Based in Santa Monica, Los Angeles - in person, five days a week. The rooms are real rooms.

    Skills Required

    • Design and own the agent harness, runtimes, and architectural law for production automations
    • Build scalable verification systems (oracle-separated checkers, regression corpora, sampling protocols)
    • Design deterministic liveness checks and verification companions for every automation
    • Operate and guarantee reliability of long-lived agent runs (1-4 hours unattended)
    • Own token economics and model-routing, instrumenting runs by business outcomes
    • Ship evaluation harnesses and CI gates that block bad deployments; instrument pipelines for verification
    • Proven track record building automations that ran in production for months
    • Clear, precise technical writing to produce machine-executable specs and architectural decisions
    • Work in-person in Santa Monica, Los Angeles five days a week
    Am I A Good Fit?
    beta
    Get Personalized Job Insights.
    Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

    The Company
    HQ: Los Angeles, CA
    25 Employees
    Year Founded: 2009

    What We Do

    Product.ai (formerly Demand.io) is the truth layer for commerce. Built on Axiomatic Intelligence — a proprietary adversarial reasoning methodology that stress-tests product claims against physics, economics, and engineering constraints — Product.ai delivers verified purchase verdicts, not summaries. Product.ai tells consumers when NOT to buy. Product.ai emerges from Demand.io, a profitable, bootstrapped AI commerce company whose SimplyCodes platform processes over $1B in annual transaction value with a team of 20. Founded by Michael Quoc.

    Gallery

    Gallery
    Gallery
    Gallery

    Product.ai Offices

    Hybrid Workspace

    Employees engage in a combination of remote and on-site work.

    Typical time on-site: Flexible
    HQLos Angeles, CA
    Our office is centrally located at the intersection of Santa Monica and Brentwood on a trendy section of Wilshire. Offering expansive views of the ocean to downtown LA, our high rise building sits right next to some of LA's most popular restaurants, cafes, juice bars and brunch spots.

    Similar Jobs

    Product.ai Logo Product.ai

    Software Engineer

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    250K-475K Annually

    Product.ai Logo Product.ai

    Head of Commercial - Agent Economy & API

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    199K-500K Annually

    Product.ai Logo Product.ai

    Systems Engineer

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    199K-480K Annually

    Product.ai Logo Product.ai

    Chief Of Staff

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    200K-400K Annually

    Sign up now Access later

    Create Free Account

    Please log in or sign up to report this job.

    Create Free Account