AI vs. AI: Why Human Phishing Training Failed
AI
financial services
August 22, 2026· 7 min read

AI vs. AI: Why Human Phishing Training Failed

Twenty years of phishing training won't work against AI-generated attacks. The real defense isn't better employees—it's automated classifiers that stop threats before humans see them.

We Spent Twenty Years Blaming Humans for Phishing. The Machines Just Fired Both Sides.

86% of people approved a dangerous command last week. Not the careless ones. Not the undertrained. The normal ones — the same ones who'd been clicking "yes" all morning and stopped reading the fine print around noon.

Anthropic buried this detail in a research update, but I keep coming back to it. Because that 86% is every phishing click rate I've ever seen. Every "verify your account" email that slipped through. Every wire transfer to the wrong account that started with someone clicking a link they shouldn't have.

For twenty years, I've lived inside the security industry's answer to this problem: train the human. Send fake phishing emails. Measure who clicks. Make the clickers watch a video. Repeat quarterly. I've run this program at multiple organizations. I've bought this program. I've watched the click rate drop from 40% to 15% to 8%.

It never hits zero. And now I know why we were asking the wrong question entirely.

The Test Was the Failure

Anthropic's experiment wasn't about phishing — it was about AI approval workflows. They slipped a dangerous command into an approval box in front of 1,053 people and watched what happened. The setup was clean: people had been clicking legitimate approvals all morning. By noon, pattern recognition had replaced careful reading. When the dangerous one appeared, 86% waved it through.

That's not a failure of training. That's a failure of design.

We built a system that requires continuous human vigilance against an infinite stream of small decisions, most of which are benign, a few of which are catastrophic. We asked people to be spam filters, one email at a time, forever. Then we acted surprised when the filter got tired.

I watched this play out in 2008 with a client in manufacturing. Their phishing training was immaculate — quarterly tests, below-industry click rates, certificates on the wall. Then someone in AP clicked a link in an email that referenced a real invoice number, from a real vendor, with the right payment cadence. The wire went out Tuesday. We found it Thursday. The click rate was 100% — the one person who needed to catch it, didn't.

The post-mortem blamed the human. But the real failure was that we'd designed a system where everything depended on that one person having a bad feeling at the right moment.

We've Seen This Movie Before

Spam used to be a human problem. I remember 2003 — you'd open Outlook and spend the first ten minutes deleting penis pills and Nigerian princes. Every company had a policy: "Don't click suspicious emails." Every employee had a story about clicking anyway.

Nobody fixed spam by training humans better. We fixed it with filters.

Bayesian algorithms. Sender reputation scores. Graylisting. SPF and DKIM. The entire infrastructure shifted from "teach the human to spot the scam" to "don't let the scam reach the human." Your inbox today isn't clean because you got smarter — it's clean because machines made the decision before you ever saw the message.

Anthropic's fix to their 86% problem was the same architectural move: stop asking. Pull the human out of the high-frequency decision loop. Put a classifier in front of the action. Let the machine say no before the approval screen ever renders.

The email world already made that bet. Most of us just didn't notice because it happened gradually, vendor by vendor, over fifteen years. The tools that quarantine the bad message before it reaches your inbox — scoring sender behavior instead of hoping you spot the typo in "Micosoft" — are the same fundamental shift: the human was never the right place to make this call.

But the Attackers Read the Same Playbook

Here's the part that should keep you up at night.

The typos are gone. The grammar's perfect. The lure references a real invoice, in your CFO's actual writing style, because a language model wrote it after reading three years of his LinkedIn posts. I'm seeing this in client assessments right now — phishing emails that pass every traditional "spot the fake" training because there's nothing to spot.

We used to tell people: "Look for the misspellings, the awkward phrasing, the generic greeting." That advice aged out six months ago. The current generation of attacks is written by the same models your marketing team uses to draft email campaigns.

Last month I reviewed an incident where an AI-generated email referenced an internal project by its correct code name, mimicked the VP's habit of ending emails with "Thoughts?", and landed in the target's inbox fourteen minutes after a real Slack conversation about that project. The recipient clicked. Of course they clicked.

One AI is writing the lure. Another is deciding whether you ever see it. The human — the one we spent two decades blaming for the click — is being walked off the field on both sides.

What Nobody Wants to Say Out Loud

I've sat in enough board meetings to know the question that's coming: "So we just let the machines handle everything?"

Not quite. But we do need to stop pretending quarterly phishing tests are a control. They're not. They're theater — security kabuki that makes us feel like we're doing something while the actual architecture of the threat has shifted underneath us.

The uncomfortable truth is that if your last line of defense against a machine-written attack is a Tuesday-afternoon employee clicking "report phishing," you don't have a defense. You have a hope. And hope isn't a strategy — it's what you call the plan after the real plan failed.

This isn't an argument for eliminating human judgment. It's an argument for repositioning it. Humans are extraordinary at high-stakes decisions with full context. We're terrible at maintaining vigilance against low-probability events in high-frequency streams. Every field that's figured this out — aviation, manufacturing, healthcare — has moved humans out of the vigilance role and into the exception-handling role.

Security is late to this realization, but we're getting there. The question isn't whether to deploy AI-powered email filters and approval classifiers. The question is what you're doing Thursday when the attacker's AI gets better than your defensive AI, and the employee who used to be your "human firewall" has been retrained to trust the filter.

What to Ask Monday Morning

I don't have a clean answer to that question. But I know the wrong answer: keep running the quarterly phishing test and hoping the click rate drops another two points.

Here's what I'd actually raise with your security team:

What percentage of our security controls depend on a human recognizing something is wrong before they click? Map those controls. If the answer is "most of them," you're running the 2015 playbook in 2025.

Where are we still asking humans to be classifiers instead of decision-makers? Approval workflows. Email triage. Incident escalation. Find the places where someone clicks "yes" forty times a day and "no" matters once a quarter.

What's the plan when the attacker's AI writes better email than our CEO? Because that's not a future-tense question anymore. I'm seeing it in assessments today.

We spent twenty years blaming humans for being human — for getting tired, for trusting patterns, for not reading the fine print on the 47th approval of the afternoon. The machines didn't fire the humans out of cruelty. They fired them because we kept assigning them an impossible job and calling it a security control.

The real question is whether we'll redesign the system before the next 86% clicks through. Or whether we'll just send them another training video and call it a day.

What do I know — I've only watched this movie three times. But the sequel always has better special effects.

Frequently asked questions

Why has phishing training been ineffective for 20 years?
Because training humans to be spam filters one email at a time doesn't work—the click rate never hits zero. Most people who clicked phishing links weren't careless; they were normal employees fatigued by decision-making, demonstrating that the problem isn't employee failure but an impossible task design.
How are AI-generated phishing attacks different from traditional ones?
Modern AI-written phishing emails have perfect grammar, no typos, and reference real details from actual company communications and executive writing styles. They're fundamentally more convincing because they're machine-generated at scale, similar to how spam evolved.
What's the better defense against AI-powered phishing?
Automated classifiers that screen emails before they reach inboxes—the same approach that defeated spam. Rather than relying on employees to spot threats, the solution is to remove humans from the decision loop and let machines evaluate sender behavior and message legitimacy.
What should organizations plan for if their main defense is employee reporting?
The post raises a critical board-level question: if your last line of defense is an employee clicking 'report phishing' on Tuesday afternoon, you have no plan for Thursday. With AI creating attacks faster than humans can respond, organizations need automated detection systems in place, not behavioral training.
Get More Insights
Join thousands of professionals getting strategic insights on blockchain and AI.

More Ai Posts

August 14, 2026

The AI Pricing Time Bomb: Your Strategy

You're paying 2% of true AI costs. Learn what happens when OpenAI and Anthropic reprice subscriptions and how to future-...

February 23, 2026

Why Solo AI Builders Are Your Market Canaries

Solo developers using AI are discovering pricing models and tools enterprises will demand in 2-3 years. Watch them to pr...

December 22, 2025

Stop Waiting for AI: Your Competition Already Started

AI disruption isn't coming tomorrow—it's happening now. While most companies debate, competitors are shipping. Here's wh...