Staff Software Engineer, AI Infrastructure

Posted Yesterday
Be an Early Applicant
Metropolitan, CA, USA
In-Office
300K-450K Annually
Senior level
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Product.ai is the truth layer for commerce.
The Role
Own the distributed serving and data infrastructure powering globally served pages and governed multi-agent build loops. Design caching, indexing, rebuild, reliability, observability, performance, and invariant-based enforcement systems. Take end-to-end responsibility for correctness, uptime, failure handling, latency budgets, and production incidents without a dedicated operations team. The role requires independently modeling complex systems, proving correctness under load, and evolving infrastructure as AI capabilities change.
Summary Generated by Built In
Design the substrate that the agent loops and the products stand on: what serves fast at scale, what an agent's output is allowed to do, and what stays up when nobody is watching it.

Product.ai is the verified truth layer for shopping, the intelligence that tells you what is actually true about a product, including when not to buy. SimplyCodes , its first proof at scale, is the code verification service, at about $22M in revenue. We are 100% founder-owned, profitable, and bootstrapped since 2009. No outside investors. No board. Fewer than twenty operators outbuilding companies 10x our size.

Why This Role Exists

Agents write much of the code here now. That only works because something else holds the line: a serving layer fast enough to answer at scale, and a body of rules an agent's output has to pass before it counts as shipped. Today those two systems live with a very small team. This seat is a third builder at the same depth, on the half of that substrate that needs the deepest systems background: the data and serving path, and the reliability floor underneath it, with no ops team to catch what falls through.

There's a sibling seat, Agent Systems, that owns the harnesses judging whether a loop's code or content output is correct. This seat owns the invariant-level gates guarding the serving path and the reliability floor underneath all of it: a different altitude of the same problem.

This is the same class of problem as the two systems above, handed to someone who owns it end to end: distributed systems and correctness under real load, with the architecture decisions actually yours to make.

The System You'll Need to Model

  • A serving layer at real scale, real latency. Hundreds of thousands of pages get rebuilt every night and served globally inside a hard latency budget. Every layer between a request and an answer, caching, indexing, the data path underneath, has to fit inside that budget, and it gets harder to hold as the page count grows.
  • A governed multi-agent build loop. Agent loops here run a full build to a shipped artifact, governed entirely by automated gates: deterministic checks, written as software, that decide what an agent's output is allowed to do, and where a judgment call needs a second, adversarial model instead of a human's eyes. Humans own the design of those gates and the failure modes they exist to catch. Your slice is the invariant layer underneath: the gates that hold when the serving path itself is what changed.
  • Correctness you can prove, not just observe. A green checkmark is a claim. The gates that decide whether millions of pages are correct, or whether an agent's output is safe to ship, need to be defended by invariant and by test, not by a demo that happened to work once.
  • Reliability with no net. There's no dedicated ops team standing behind this. When something breaks at 3am, the person who designed the system is the one who explains why it broke and what invariant would have caught it sooner. Correctness and uptime are your problem from design through the pager.
  • A partial rebuild is still a live system. When a nightly rebuild fails partway through, the cache underneath has to decide what's stale, what's safe to keep serving, and what the reliability floor needs to catch before a customer sees a wrong answer. That interaction, not the rebuild itself, is where the hard bugs live.
  • The ground moves under you. Every model generation changes what an agent loop can be trusted to do unsupervised, which means the gates and the serving layer both get rebuilt while they're running, not designed once and left alone.


If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.

What You Will Own

  • The serving and data path. The infrastructure that rebuilds a huge page surface nightly and answers requests globally inside a hard latency budget. You own the architecture of what gets cached, what gets indexed, and what gets recomputed, and you defend the number when it's under pressure, not just when it's easy.
  • A slice of the enforcement layer, at the invariant level. The gates that decide whether the serving and data path holds under the load and the changes an agent introduces. Judging a single agent's code or content output belongs to the neighboring Agent Systems seat; you design these as software, as invariants and tests, not as prose policy someone reads and forgets.
  • The reliability floor. Uptime, correctness, and failure modes for systems that run unattended. You build the observability that tells you a system is wrong before a customer does, and you're the person paged when it doesn't.
  • The performance envelope. Where the system's real bottleneck is, and what it costs in latency or dollars to move it. You make that tradeoff explicit and defensible instead of folklore.
  • Your seat charter. Within your first quarter you co-sign a charter for this seat. It names one machine-checkable number that proves the seat is working, and a written split of what you decide freely versus what you bring to the founder to decide.


You will use the craft you already own (distributed systems, performance engineering, systems design defended under questioning) and grow into the layer above: designing the governance a multi-agent build loop runs inside, and the invariants that let autonomy scale without a human checking every step.

Who You Are

You reason about a system from its failure modes first. Handed an unfamiliar service under real load, you can name the three ways it breaks before you've read every line, and you can tell the difference between a system that's correct and one that merely hasn't failed yet. You form a working model fast, notice when it's wrong, and update. You write clearly, because clear writing is evidence of clear thought, and because the invariant you can't state in one sentence is one you haven't actually proven.

You treat agents as a tool you verify, never as an oracle you trust. You can still do the work by hand: trace a latency regression to its root cause, design a cache invalidation scheme, prove an invariant holds under a race you constructed yourself. That mastery is what lets you look at code an agent produced and say why it's right, or refuse to ship it.

You've probably built or rebuilt a system that serves real traffic under a hard latency budget, designed the correctness checks a team came to depend on without asking whether they still held, or debugged a production failure that only showed up under load nobody had tested for. Adjacent roads count fully: high-throughput data platforms, real-time serving infra, distributed systems under load anywhere it was unforgiving. We care about the artifact and the reasoning more than where you built it.

Who this isn't for. This is wrong if your instinct with a hard system is to describe it at a high level rather than trace it down to the invariant that actually holds. It's wrong if you want a fixed lane and a spec handed to you; this seat models the system and proposes the fix. It's wrong if you want the staff title for the resume line rather than the invariant work underneath the latency budget and the 3am pager that earns it. It's wrong if you're comfortable shipping what an agent produced without being able to defend why it's correct. And it's wrong if your instinct at 3am is to manage the story about what broke instead of naming the invariant that would have caught it sooner. You'll be happiest here if the phrase "no ops team" reads as ownership, not as a gap someone else should fill.

How We Evaluate

We don't run traditional engineering interviews.

  • Async video screen. About 15 minutes, on your own time. It replaces the recruiter screen. We want to see how you think, not how you present.
  • Calls with company stakeholders. Short conversations with the team you'd build beside.
  • Conversation with the founder. How you model the system above, where you push back, and whether you can defend the argument live.
  • Paid work trial. Real work, in our real environment, on the systems above. We watch how you get grounded, whether you write the invariant before the build, and whether your self-assessment is honest.


  • If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.

    Compensation & Ownership

    Total first-year comp: $450,000 to $550,000 (base + performance-based ownership and profit-share programs). Base: $300,000 to $380,000, top of market for staff-level systems engineering. Eligibility for the company's ownership and profit-share programs; grants are performance-based, terms discussed at the offer stage.

    100% premium coverage for you and your family. The token budget is effectively unlimited, steered by return, never capped.

    Based in Santa Monica, Los Angeles, in person, five days a week. Relocation support available for the right builder.

    Skills Required

    • Deep experience with distributed systems
    • Experience designing or operating systems serving real production traffic under hard latency budgets
    • Performance engineering and systems design expertise
    • Ability to design cache invalidation, data serving, indexing, and rebuild systems
    • Experience defining correctness checks, invariants, and automated tests
    • Experience debugging production failures under high or unexpected load
    • Ability to own reliability, uptime, observability, and on-call response end to end
    • Ability to reason from system failure modes and defend architecture decisions
    • Clear technical writing and communication
    • Experience with high-throughput data platforms or real-time serving infrastructure
    • Experience designing governance or enforcement systems for multi-agent build loops

    Product.ai Compensation & Benefits Highlights

    • Healthcare Strength — Health coverage is described as 100% employer-paid for employees and families across medical, dental, and vision. Multiple public listings also cite immediate eligibility on the start date.
    • Leave & Time Off Breadth — Time off policies include unlimited PTO, with the company explicitly stating it expects employees to use it. Public materials also reference paid holidays and sick days.
    • Equity Value & Accessibility — Compensation includes profits-interest units and profit sharing, with an annual tender enabling employees to sell a portion of vested units back to the company. This structure is positioned to provide ownership-style financial upside as part of total rewards.

    Product.ai Insights

    Am I A Good Fit?
    beta
    Get Personalized Job Insights.
    Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

    The Company
    HQ: Los Angeles, CA
    25 Employees
    Year Founded: 2009

    What We Do

    Product.ai (formerly Demand.io) is the truth layer for commerce. Built on Axiomatic Intelligence — a proprietary adversarial reasoning methodology that stress-tests product claims against physics, economics, and engineering constraints — Product.ai delivers verified purchase verdicts, not summaries. Product.ai tells consumers when NOT to buy. Product.ai emerges from Demand.io, a profitable, bootstrapped AI commerce company whose SimplyCodes platform processes over $1B in annual transaction value with a team of 20. Founded by Michael Quoc.

    Gallery

    Gallery
    Gallery
    Gallery

    Product.ai Offices

    Hybrid Workspace

    Employees engage in a combination of remote and on-site work.

    Typical time on-site: Flexible
    HQLos Angeles, CA
    Our office is centrally located at the intersection of Santa Monica and Brentwood on a trendy section of Wilshire. Offering expansive views of the ocean to downtown LA, our high rise building sits right next to some of LA's most popular restaurants, cafes, juice bars and brunch spots.

    Similar Jobs

    Product.ai Logo Product.ai

    Product Engineer

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    200K-425K Annually

    Product.ai Logo Product.ai

    Design Engineer

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    130K-230K Annually

    Product.ai Logo Product.ai

    Head Of Product

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    250K-475K Annually

    Product.ai Logo Product.ai

    Technical Product Manager

    Artificial Intelligence • Big Data • Consumer Web • eCommerce
    In-Office
    Metropolitan, CA, USA
    25 Employees
    200K-425K Annually

    Sign up now Access later

    Create Free Account

    Please log in or sign up to report this job.

    Create Free Account