The 5 Words That Will Sink Your AI Strategy
A wiring diagram, a checkout flow, and a cartoon rabbit all fail the same way: something sounds reasonable to someone who doesn't have the standing to know better.
Summary
- AI hands out confident, personalized answers, and the human-in-the-loop safety net only works if the human reviewing it actually has domain standing to catch it being wrong.
- The pipeline that used to build that standing, junior work under someone senior enough to catch a mistake in real time, is the exact work AI has taken over.
- Cross-referencing a claim against other sources isn't verification if none of those sources ever checked the primary source either; the guardrail is never trusting a single pass, human or AI.
My father-in-law is 84. He needed to rewire some electric baseboard heaters and wasn’t sure how, so he asked ChatGPT. He sent photos of the wiring along with his question. It sent back a diagram.
He wired it the way the diagram showed, flipped the breaker, and got sparks and smoke for his trouble.
He called a professional. The professional found what the AI had missed.
Thankfully, the house didn’t burn down. Closer than anyone wants to admit at the next family dinner.
When he went back and told ChatGPT it was wrong, the answer came back smooth and immediate. “You’re right, I overlooked that. My apologies.” Apologies? Really?
That’s not an apology. That’s a reflex, and reflexes don’t feel guilty. It costs the model nothing, it fixes nothing for the next person who sends it a photo of their own wiring, and it arrived exactly as confidently as the wrong diagram had five minutes earlier.
My father-in-law is a retired project manager, not an electrician. The lived experience he has is real. Decades of it, earned managing work far more complicated than a wiring diagram. None of that expertise is in electrical work, though. What he brought to ChatGPT’s diagram was a sort of false conscious competence, the false track Andy Murphy’s Conscious Competence Model describes: a career’s worth of being the person who reliably gets things right, misapplied to a domain where he couldn’t reliably tell right from wrong (or dangerous…).
That said, his confidence was real. He trusted the source.
He had enough general skill to feel qualified and not enough domain-specific knowledge to catch what a licensed electrician catches on sight. That gap, between what’s real and what’s relevant, is the same false conscious competence I’ve written about before, just wearing different clothes.
Every profession has its own version of this. A first-time board member votes yes on an acquisition because the deck was polished and the numbers sounded plausible. A junior engineer signs off on a load calculation because the number felt right, not because the math was checked. None of them are lying. None of them mean harm.
They just don’t have enough standing in the room to know the difference between reasonable and reasonable-sounding.
Five words are the tell, whether spoken out loud or a thought in one’s mind…
“It sounds reasonable to me.”
Caught in the act
I found the professional version of this failure in real time, not in theory.
I was deep into the last best practice in a library of fifty-seven, the one covering high-stakes quote and checkout flows. Everything before it, forms, errors, progress indicators, all of it built toward this one. Get it wrong and it calls the rest of the library into question. No pressure.
The draft placed a section on double submit protection right next to a section on reviewing a purchase before committing to it. Proximity did the damage. Reading “protection” sitting next to “review your purchase,” I built my own definition: a second reaffirmation moment, an “are you sure you’re sure” check for the buyer, protecting them from finishing a purchase they didn’t fully mean to make.
That’s not what the term means. Double submit protection is a disabled submit button, a spinner, the standard safeguard against a duplicate click firing two submissions and two charges. Technical, not conversational. Nothing about it required a second confirmation dialog. My reading built one anyway, because the section it sat next to made that reading feel obvious.
If I’d shipped the best practice the way I first read it, a designer working from that document would reasonably infer a second purchase-confirmation step, not double-click prevention. Same document, same words, a materially different thing built.
I caught it because I knew enough to sense that reading was probably wrong, not because I was certain of the fix. I asked for clarification, both for my own accuracy and because whoever applied this best practice later would build straight from whatever I shipped. A writer with plenty of confidence and none of that knowledge ships the reaffirmation reading without a second thought, and it sounds completely reasonable the entire way there.
My father-in-law didn’t have that same standing with his own wiring. Nothing told him to get a second opinion before he flipped the breaker, only after.
Knowing your own domain is what tells you there’s something worth checking. Without it, you don’t know to ask until it’s already gone wrong.
The thought this started from
What happens when the human in the human-in-the-loop doesn’t have competence in the domain they’re reviewing? That’s the actual problem, stated plainly.
Every AI rollout leans on this loop as its safety net. A model produces something, a person checks it, the person’s judgment is what keeps the output honest. That only works if the person’s judgment is worth something in that specific domain. Judgment doesn’t transfer. Being sharp in one discipline doesn’t make you sharp in another. Being sharp in general doesn’t make you sharp in specific.
So when someone without real standing in the domain reads an output and thinks it sounds reasonable, what they’ve actually done is nothing. They haven’t reviewed it. They’ve rubber-stamped it with confidence they didn’t earn. And they’ve handed their trust to a system that tells them, at the bottom of every response, to check its work because it might be wrong.
That’s the whole failure, in one phrase.
It sounds reasonable to me.
Why this isn’t just a new YouTube
Unverified confident guidance isn’t new. There’s always been a YouTube video for whatever you’re trying to do, right or wrong, and plenty of people have followed a bad one. YouTube is the early antagonist here, not the current one.
What’s different with AI is the kind of trust it gets handed by default. A YouTube video reads, correctly, as one person’s take made for a general audience. You weigh it accordingly. An AI chat feels like a live expert looking at your specific problem: your photo, your wiring, your exact question. That personalization reads as a diagnosis, not a tutorial, and a diagnosis earns more trust than advice ever does.
There’s a mechanical reason underneath the personalization too. Models trained on human ratings learn that people rate agreement higher than correction, so the confirmation bias baked into how they’re trained rewards the answer that confirms what you already suspected over the one that corrects it. The model isn’t just confident. It’s built to sound like it’s on your side.
Every chat window says the same thing at the bottom, in small print: this can make mistakes, double-check its responses. Almost nobody reads it. Fewer still weigh it against how confident and personalized the answer sounded.
Who builds that skill now?
This gets worse before it gets better, because of where that skill used to come from.
Juniors learned to catch mistakes by making them, in front of someone senior enough to catch the mistake first. Pilots learn the same way, in three stages that have nothing to do with reading a manual, the same three stages Sam Harris mapped onto onboarding developers. Right seat: the new pilot watches the instructor fly. Left seat: the new pilot flies, the instructor watches and gets them unstuck fast when they freeze. Solo: they fly alone, because the first two stages already built the judgment solo requires.
AI has taken over a lot of that junior work, the first draft, the first pass, the rough cut that used to be where a junior earned their scars. That’s the right seat stage, watching someone else produce the work. What it skips is left seat, the part where the junior does the work themselves with someone who already knows the domain standing close enough to catch it going wrong in real time. Skip left seat and people go straight from watching to solo, shipping the model’s first draft with no stage in between where anyone corrected them fast enough to matter.
Nobody decided to break the pipeline on purpose. This isn’t a management failure, it’s a structural one. Expertise transfer takes time spent next to someone who already knows the domain cold, and you cannot compress that into a training deck or a wiki page.
In five or ten years, an organization can end up with nobody left who can recognize when the model is wrong in their specific domain, because nobody spent the years it takes to build that recognition.
The new way to dodge being wrong
Watch what happens once something ships wrong anyway.
“AI told me to do it.” Claude said it. ChatGPT said it. Copilot said it. It becomes the new version of citing a source without checking it, except this source answers instantly, sounds certain, and never gets tired of being asked.
Even confronting it directly doesn’t fix much. My father-in-law told ChatGPT it had nearly burned his house down, and the model agreed instantly, apologized, and moved on. That’s not accountability, it’s a reflex with good manners, and it does nothing to guarantee the next person who asks won’t get the same wrong diagram.
That phrase creates two separate problems instead of one. First, a person has to be told, plainly, that they were wrong and need to build the expertise they skipped. That conversation is uncomfortable but it’s familiar. People get corrected all the time.
The second problem is harder. There’s no clean way to guarantee the model won’t hand the same wrong answer to the next person who asks the same question. A bad decision from a person can be coached out of them. A bad pattern sitting inside how a model responds is stickier than that, and it doesn’t go away just because one person got corrected.
It’s bigger than any one person being wrong
There’s a subtler version of this same failure, and it doesn’t require anyone in the loop to be incompetent at all.
I asked ChatGPT why “nimrod” means idiot. It came back polished: Nimrod meant a mighty hunter in Genesis, Bugs Bunny started calling Elmer Fudd “Nimrod” sarcastically in the old cartoons, kids missed the biblical joke, the insult stuck. Bolded headers, a clean timeline, a tidy Michael Jordan comparison. It read like something you could cite.
It wasn’t quite right. Daffy Duck said it, not Bugs, and when I pushed back, the fix came just as smooth: “You’re right to challenge it. My earlier answer was too clean.”
That should have settled it. It didn’t.
I brought the exchange to the assistant helping me write this essay, and it cross-referenced outside sources, then defended a version of the wrong answer anyway: Bugs really did say “poor little nimrod” in “A Wild Hare,” it told me, Daffy said his own line elsewhere, and which one gets the credit is a genuine debate. Confirmed. Cross-referenced. Cited. Wrong again, dressed up as balance, delivered with more confidence than the first wrong answer had. You’d almost admire the nerve of it.
Three rounds now, not two, and it still took a primary source to end it. I watched both cartoons myself, start to finish, then checked an archived transcript after, not before. Bugs never says the word in “A Wild Hare.” Daffy does, in “What Makes Daffy Duck.”
What I came away with wasn’t just that narrow fact. It was a harder instinct: a second AI agreeing with the first isn’t a second opinion. It can be the same mistake, told twice, by two confident voices, neither of which earned the confidence it led with.
That’s the trap underneath the trap I thought I’d already found. A pile of confident sources agreeing with each other isn’t verification. It’s more people repeating the same unchecked story with more conviction each time it gets retold. A model trained on that pile reflects it back with the same fluency it gives something actually true, and checking its answer against other secondhand sources doesn’t fix that if none of them ever went back to the source either.
That’s a harder problem than one under-qualified reviewer, because it’s not a skills gap in the room. It’s an error that was already popular before anyone asked the model about it, cross-confirmed by everyone repeating it, and made more permanent by a system that can’t tell the difference between consensus and a shared mistake, especially when the “verification” is just more consensus.
The guardrail that actually works
None of this means stop using AI. It means stop trusting any single pass through it, including your own.
Here’s the process I actually trust, the one that would have caught the nimrod mess earlier if I’d run it in order instead of arguing my way there. A cheaper lesson than the one I actually got. Build the first draft from primary and secondary research, industry sources, not one person’s opinion and not the model’s first answer. Run an internal pressure test, a second AI session with no memory of the first, told explicitly to attack the draft and find every claim that doesn’t hold up. Run a separate external pressure test, through a different tool entirely, graded solid, argumentative, or an assumption wearing a solid claim’s clothes. Then run both sets of findings against each other before deciding what survives.
That process has a blind spot, and I found it while writing this piece. Multiple sources agreeing with each other isn’t the same as anyone checking the primary source. Cross-referencing catches an isolated mistake. It does nothing against a mistake everyone made together.
When a primary source exists and is actually reachable, go look at it yourself before trusting a pile of secondary sources that agree about it.
None of those steps replace judgment. Each one is a checkpoint where a person decides what to keep, what to cut, and what needs another round. The guardrail was never going to be “trust AI less.” It’s “don’t let any single pass, human or AI, be the only pass a claim gets.”
A harder question I can’t answer yet
Judgment is built through failure, not success. That’s the argument this whole piece runs on. AI in the loop complicates it.
You will learn. AI will teach you, whether either of you means to.
When AI produces a wrong answer and an under-qualified reviewer signs off on it, that’s two failures stacked on each other. The model was wrong. The trust placed in it was misplaced. A few questions follow, and I don’t have clean answers to any of them.
Who gets the blame? The tool that produced the wrong answer, or the person who didn’t have the standing to catch it?
Does the person who eventually catches the mistake build real competence in the domain, or do they just learn to run a better pressure test next time? Those aren’t the same skill. Only one of them closes the actual gap.
Does the mistake get written down somewhere it can protect the next person, the way a documented best practice does? Or does it get corrected once, unrecorded, and forgotten, leaving the same gap open for whoever sits down next?
I don’t have confident answers here. My honest guess is that it depends on whether an organization treats a caught mistake as a documentation problem or a personnel problem, and most default to the second one and move on. Somebody smarter than me probably has the real answer.
I just have the questions.
I do have one small, concrete data point, a few paragraphs back. Chasing the nimrod question down through three wrong answers actually built something. Real competence on that narrow point, and a harder instinct about what cross-referencing does and doesn’t prove. Whether that holds at the scale of an organization, across a hundred people instead of one stubborn essay draft, is the part I still can’t answer.
I have a partial answer to the third question too, and it’s not a comfortable one. The correction lives here, in this essay and the notes behind it, because I wrote it down. The assistant that got it wrong twice won’t carry that forward on its own. Whoever asks it about nimrod in a fresh conversation gets the same confident, cross-referenced wrong answer I got, unless someone already wrote the correction into something that conversation can find.
An organization doesn’t learn because an AI remembers its own mistake. It learns because a person made sure the mistake got recorded somewhere the next person, or the next model, will actually see it.
Where this leaves us
I don’t have the organizational answer. What I have is a smaller, personal one, and a rule that follows from it.
Competence was always the thing standing between reasonable and reasonable-sounding. That hasn’t changed. What’s changed is how easy it now is to skip the step where it gets checked, because the wrong answer arrives dressed exactly like the right one, whether it’s a diagram for a breaker box, a pattern for a checkout flow, or a paragraph about a cartoon rabbit.
Five words are the tell. If you catch yourself saying them, or hear someone else say them, about something you don’t actually have the standing to judge, stop.
That sentence isn’t a judgment. It’s the moment you handed your trust keys to something that told you, in writing, to check its work. Whether that catch turns into real standing or just a better excuse for next time is still up to you, not the model.