Turns Out ‘Stop Hiring Humans’ Was Terrible Advice

Businesses buying enterprise AI still can’t prove whose judgment they are buying.

Written by Terence Kwok
Published on Sep. 03, 2026
A marble statue of a Greek philosopher
Image: Shutterstock / Built In
Brand Studio Logo
REVIEWED BY
Summary: Companies attempting full AI automation are facing quality issues and bringing human workers back into workflows. Rather than replacing headcount, enterprise AI succeeds through human-AI collaboration backed by verifiable, attributable expert credentials to ensure judgment and accuracy at scale.

Artisan bought billboards across San Francisco in 2024 that read “Stop Hiring Humans.” The promise underneath was payroll. Software does the work; headcount goes down. Two years later, the replacement thesis looks far less convincing. Companies that pushed hardest toward automation have run into quality problems, and some have brought people back into workflows they expected AI to own.

In 2024, Klarna said its AI assistant was doing the work of 700 agents. By 2025, CEO Sebastian Siemiatkowski was publicly conceding that the cost focus had driven down quality, and Klarna began recruiting human agents again.

A convincing output no longer proves who’s standing behind it or whether that person is qualified. A well-formatted answer about drug interactions or lease termination used to imply that someone qualified wrote it. That inference is dead. Buyers of enterprise AI now need to know whose judgment is embedded in a system and how much of it can be verified.

Why Does AI Not Reduce Headcount?

Aggressive AI headcount replacement is leading to quality issues, forcing enterprise buyers to shift from full automation to human-AI collaboration. True enterprise AI success depends on verifiable expert authority — tracking whose credentials, jurisdiction and tying review date back an output — to ensure accuracy, auditability and fair economic attribution for human experts.

More From Terence KwokAssume Everyone Is AI Until Proven Otherwise

 

Replacement Was the Wrong Benchmark

Headcount gives executives an easy AI success metric because it appears directly in a budget. Skilled judgment is harder to measure. Experienced employees know where procedures fail and which exceptions matter, and they recognize when a decision needs escalation. Little of it is written down anywhere.

Anthropic’s March 2026 Economic Index found that users who had signed up at least six months earlier were more likely to work with Claude in several steps, revising along the way, instead of handing over a task in one instruction. Their conversations succeeded 10 percent more often. After controlling for model, language, task and country, 4 percentage points of that gap persisted. The adjusted effect is modest, but it supports a narrower point about experienced users getting more from iterative collaboration.

Aggressive replacement removes the people on whom automation depends. Organizations keep the workflows and documents after key employees leave, but not the reasoning behind them. AI can execute a process at enormous speed. But when a process reaches its limits, someone has to decide what to do next, and that decision must come from a source the business can vouch for.

A polished answer proves nothing about provenance. Models generate plausible legal or technical guidance from strong source material. They also generate it in the same tone from a statistical guess. The output looks identical either way.

Detection tools won’t fix this problem. Every model generation writes more like a person than the last, so the detectors lose ground over time. The quality check has to happen earlier, at the point where the knowledge enters the system. A business must prove that a real, credentialed professional produced the information and that the software has permission to use it.

 

Expertise Needs Verifiable Claims

Enterprise AI should carry evidence about the qualifications behind its outputs. Imagine an AI agent giving a company guidance on a regulated filing. Before acting, the buyer should be able to verify that the agent speaks for a named specialist, that the specialist holds the relevant credential for that jurisdiction, and that the underlying guidance was reviewed on a known date.

An agent booking a meeting needs the first of those. An agent drafting a dosage recommendation needs all four, plus a check performed at the moment it answers. The greater the cost of a wrong answer, the more friction the system needs.

Claim-centric trust means verifying the specific facts a decision depends on, such as credential status, jurisdiction and review date. That narrower approach matters for privacy because the system can prove professional authority without collecting a complete identity profile. A credential check does not require a person’s passport number and home address.

Selective proof of the claims behind an expert’s authority gives businesses the assurance they need while shrinking their data footprint. A medical agent, for example, could prove that the clinician’s professional license is active and their specialty matches the question being asked. The system never needs the clinician’s birth date or home address. Holding less personal data also gives a potential attacker less information to steal.

Verification also creates an audit trail. A business can record which expert credentials and review dates supported an answer, then route the case back to a person when those verified claims no longer cover the question being asked. The system can scale routine judgments while preserving a record of whose expertise supported them.

 

Scalable Expertise Needs an Economic Model

McKinsey has identified scalable, encoded expertise as a foundation for enterprise-ready agentic AI. The business case extends beyond internal productivity. AI distributes scarce professional knowledge far beyond what one person can handle in consultations or reviews.

Senior engineers could encode review standards for hundreds of teams, only needing to step back in when unusual cases arise. A compliance specialist covering multiple jurisdictions could keep the same verified guidance available on demand, instead of restating that guidance for each case. The subject-matter expert still anchors the underlying judgment even when software handles the delivery.

Scale also changes who captures the economic value created when expert knowledge is reused. Most engineers, clinicians, compliance specialists and other professionals are paid by salary or by the hour. Once their knowledge is packaged into an AI system and reused thousands of times, that pay structure no longer describes what they actually contribute.

A compliance specialist could license a verified body of cross-border guidance to an AI agent, define which jurisdictions and use cases it covers, and receive compensation when businesses use that expertise. The buyer gets a traceable source and clear usage limits. The specialist can reach far more organizations without repeating the same consultation from scratch.

Weak attribution of expert knowledge creates a fragile economic model. When an AI system reuses a professional’s judgment without tracking whose knowledge it used, the expert can lose the link between contribution and compensation. An AI economy built on reusable expertise needs that link if it expects experts to keep contributing new, current judgment.

How Is AI Reshaping Work?Is AI Creating a New Class of Knowledge Workers?

 

Expertise, Verified at Scale

The infamous Artisan billboards argued for measuring AI progress in headcount removed from the org chart. Enterprise buyers are now asking a duller question — which credential, whose review and when was it checked?

For teams building enterprise AI around verified human expertise, the implementation priorities are straightforward. Match the proof to the consequence, collect only what the check requires, attach attribution to the knowledge itself and decide in advance which cases go back to a person.

AI is good at getting one expert’s judgment to more people than that expert could ever reach. The harder problem is making that expertise verifiable, attributable and economically sustainable as it scales. Buyers need to know whose judgment they are using, when it was validated and where it applies. Likewise, experts need a way to retain attribution and participate in the value their knowledge creates.

Explore Job Matches.