Why AI Safety Needs Humans (Even When They Fail)
AI
financial services
August 15, 2026· 8 min read

Why AI Safety Needs Humans (Even When They Fail)

Anthropic's phishing test revealed humans fail 86% of approval checks—but replacing them with AI classifiers that miss 11% of threats isn't the answer. The real safety control is intentional human oversight when it matters.

We Just Fired the Humans. Now What?

Anthropic ran a phishing test on 1,053 of its own users last month. The failure rate: 86.4%.

Not 86.4% of novices clicking malicious email links. 86.4% of paid participants approving dangerous commands inside Claude Code — the kind that delete files or exfiltrate data — because those commands appeared in the forty-first approval box they'd seen that hour.

The fortieth box was safe. So was the thirty-ninth. The forty-first looked identical.

So Anthropic did what any rational engineering team would do: they fired the humans and replaced them with a classifier.

The Experiment Nobody Wanted to Run

Here's how it worked. Claude Code is an AI coding assistant that runs terminal commands. Every risky command triggers an approval prompt: "Run this? Yes / No." Anthropic paid people to work in the environment, let them build trust with the system over dozens of legitimate prompts, then quietly injected a malicious command mid-session.

86.4% clicked yes without reading it.

The new system flips the script. Instead of asking tired humans to spot danger on their forty-first decision of the hour, a separate AI classifier scans each command before execution and blocks anything suspicious. On the same test dataset, it caught 89% of the dangerous commands that humans approved. The machine blocked 800 commands people waved through. Humans blocked 6 the machine missed.

Read that gap again. The human wasn't providing oversight. They were providing the illusion of it.

The Pattern We Keep Replaying

I've watched this movie before — not with AI, but with every technology layer that got too fast for human reaction time.

In the 1980s, NYSE floor traders were the human safety layer. They'd eyeball order flow, catch the fat-finger trades, pull the brake when something looked wrong. Then electronic trading arrived. Orders moved in milliseconds. A human scanning a screen couldn't keep up, so we automated the scanning. Circuit breakers replaced gut instinct. Algorithms replaced the guy who "just knew" when volume felt off.

Did we eliminate errors? No. We just shifted to a new class of failure — flash crashes triggered by algorithms nobody fully understands, moving faster than any kill switch can catch.

Anthropic's experiment is the same pattern playing out one layer up the stack. We built AI assistants that operate faster than human comprehension, then asked humans to approve each step. It lasted exactly as long as floor traders lasted once HFT arrived.

Approval Theater vs. Actual Security

Here's the part that keeps me up at night, and it should keep you up too if your firm is evaluating AI coding tools, autonomous agents, or anything that touches production systems.

The old control was a bored human clicking yes. The new one is a classifier that scored 89% on a test it knew was coming.

That 11% gap isn't a rounding error when you're talking about systems with access to financial records, customer data, or production infrastructure. It's a door. And unlike the tired human who might catch something "weird" on their forty-second prompt after coffee kicks in, the classifier has no gut instinct, no ability to pause and say "wait, this feels off."

Point a real attacker at it — a poisoned dependency, a carefully crafted malicious repo, a command designed specifically to fool the classifier — and 89% accuracy becomes a penetration testing roadmap.

Anthropic knows this. Their documentation still tells users to manually review anything touching production systems. Which brings us to the uncomfortable question nobody wants to answer.

The Question Your Security Team Needs to Sit With

If the classifier is good enough to replace humans in the approval loop, why does Anthropic still recommend human review for production?

And if it's not good enough to trust unsupervised, why did we automate the approval in the first place?

We didn't solve approval theater. We automated the usher.

The honest answer — and I say this as someone who advises clients on AI risk frameworks — is that we're making a calculated bet. A classifier that never gets tired, never loses focus on the forty-first prompt, never clicks yes because it's 4:47 PM on Friday is better than the exhausted human it replaced. But "better than exhausted" isn't the same as "safe."

This is the gap where my clients get stuck. Your CISO wants to know: would you let this run unsupervised on production systems? Or does the human come back the moment real money is on the line?

What Fatigue Looks Like at Scale

I need you to understand how we got here, because it wasn't one bad decision — it was a thousand small ones, all responding to the same problem.

Humans are terrible at repetitive security decisions. We've known this for years:

  • MFA fatigue attacks work because people approve the seventh push notification without reading it

  • Email allowlists exist because people click links in messages that "look fine"

  • Certificate warnings get ignored because users see them dozens of times for legitimate reasons

Every "click to approve" prompt in your security stack exists because someone, somewhere, approved something they shouldn't have while tired. We kept adding human checkpoints as if more approval boxes would make humans more careful. It didn't. It just made them more tired.

The AI coding assistant problem is the same failure mode, accelerated. Claude Code doesn't generate seven approval prompts a day. It generates seventy. The fortieth looks like the thirty-ninth. The system trains you to click yes.

The Trade We're Actually Making

Let me be clear: I'd take the trade Anthropic made.

A classifier that evaluates every command with fresh eyes beats a human who stopped reading prompts six hours ago. But I won't pretend 89% on a controlled test equals "solved." It equals "better than the broken thing we had before."

The real question isn't whether AI can outperform exhausted humans at repetitive security tasks. It can, and it will, and we should let it. The question is what happens when we forget that 89% isn't 100%, and we start treating the automated control like it's actually autonomous.

Because here's what I keep seeing in client environments: we automate a security control, it works pretty well, we gradually increase our trust in it, and then one day we realize nobody's actually checking whether it's still working. The monitoring we said we'd keep in place gets deprioritized. The human review we promised becomes a quarterly audit. The quarterly audit becomes annual.

The control doesn't fail all at once. We just slowly stop watching it.

What to Actually Do Monday Morning

If you're evaluating AI coding assistants, autonomous agents, or any system that makes security decisions faster than humans can review them, here's what to ask your security team:

  1. What's our false negative rate, and how do we measure it? "89% on the vendor's test" isn't an answer. What's our ongoing validation process?

  2. What class of attacks is the classifier blind to? Every ML model has known failure modes. Which ones can't this system catch by design?

  3. Where does the human come back into the loop? Not in theory — in practice. What's the threshold where we stop trusting automation and require manual review?

  4. What happens when the classifier is wrong? Do we have detective controls downstream? Blast radius limits? A way to catch the 11% that got through?

Don't ask whether AI is better than humans at repetitive security decisions. It is. Ask what happens when the AI is wrong and nobody's watching.

The Uncomfortable Middle

We're entering a phase where AI can outperform humans at specific, narrow tasks — including security tasks humans were never good at in the first place. That's real progress. I'm not here to pretend otherwise.

But progress doesn't mean solved. The old control was a bored human clicking yes without reading. The new control is a classifier that's right 89% of the time on a test it knew was coming. We traded one class of failure for another.

The companies that get this right won't be the ones who eliminate humans from the loop entirely. They'll be the ones who figured out exactly which decisions to automate, which decisions still need human judgment, and — crucially — how to tell the difference.

Anthropic ran this test on their own product and published the results. That's the kind of intellectual honesty we need more of. They didn't ship a press release saying "AI solves security." They said "here's the failure mode, here's our mitigation, here's where you still need to be careful."

That's the model. Not "automate everything" or "trust nothing." But "automate the things humans are provably bad at, measure the failure modes obsessively, and stay uncomfortably aware of the gap between 89% and safe."

The classifier is better than the tired human. But if the only thing standing between an autonomous agent and your production data is one probability check you're not actively validating, that's not a safety control.

It's a coin you flip a thousand times a day and never watch land.

Frequently asked questions

What did Anthropic's phishing test reveal about human approval processes?
Anthropic tested 1,053 users in Claude Code by hiding a dangerous command in an approval box after users had already approved 40+ similar commands. 86.4% clicked yes on the risky command, not from carelessness but from decision fatigue—they'd already approved the same-looking prompt dozens of times.
How did Anthropic replace the human approval process?
New Claude Code sessions now default to 'auto mode,' where a machine classifier automatically reads each command and blocks dangerous ones. On the same test, the classifier caught 89% of the risky commands humans had approved, while humans only blocked 6 commands the classifier missed.
Why is an 89% effective classifier not sufficient as a safety control?
An 11% miss rate means a door remains open for real attackers. The post argues this isn't a safety control—it's a probability check you flip thousands of times daily. Genuine safety requires humans back in the loop for production-level decisions where real money is at stake, even if they're imperfect.
What's the difference between automating approval and building real safety?
Automating away tired humans improves baseline safety compared to decision-fatigued approval clicks, but replacing human judgment entirely with a classifier is 'approval theater'—it looks safer without being safe. Real safety means intentional human oversight returns when the stakes are highest.
Get More Insights
Join thousands of professionals getting strategic insights on blockchain and AI.

More Ai Posts

August 14, 2026

The AI Pricing Time Bomb: Your Strategy

You're paying 2% of true AI costs. Learn what happens when OpenAI and Anthropic reprice subscriptions and how to future-...

February 23, 2026

Why Solo AI Builders Are Your Market Canaries

Solo developers using AI are discovering pricing models and tools enterprises will demand in 2-3 years. Watch them to pr...

December 22, 2025

Stop Waiting for AI: Your Competition Already Started

AI disruption isn't coming tomorrow—it's happening now. While most companies debate, competitors are shipping. Here's wh...