When I started using AI tools seriously, they impressed me the way a car impresses you on a test drive. Everything felt effortless. The work came back fast and polished, and I handed more and more of it over to AI because those early results had earned it.
Then I took it onto the actual highway, under real load, and started noticing things.
My confidence in these tools has been on a slow correction ever since — not a collapse, a correction. And the longer I sat with it, the less it felt like the tool had let me down and the more it felt like my expectations had been set in a showroom. The version that impressed me early and the version I work with every day are the same thing. What changed was me.
The moment that crystallized it was almost embarrassingly small. In late May, I asked ChatGPT, Gemini, and Claude how many days of the week have a “d” in them. All three got it wrong. When I asked Claude again in a fresh conversation, it got it wrong in a different way. The answer is seven — every day ends in “-day.” None of them got it right, and each delivered its wrong answer with complete confidence.
That’s the part that stays with me. Not that it missed something trivial, but that it sounded exactly as sure when it was wrong as when it was right — and that I’d been extending that same confidence to questions I couldn’t check as easily.
Here’s where I landed, and it’s why I’m writing this. A car doesn’t become a different car because you’re frustrated with it on the highway. These tools aren’t going to learn from my corrections in any lasting way. If the relationship is going to work, the adjusting is mine to do — not theirs. That sounds obvious written down. It did not feel obvious while I was living through it.
So rather than just write about that myself, I tried something. I asked Claude to write about its own unreliability — in its own voice, and to be honest about what it means for whether anyone should trust it. One rule: no reassuring bow at the end, no turning the confession into a pitch.
What came back is more candid than I expected. It admits, plainly, that it can’t be taught in any lasting way, that it isn’t even consistent with itself from one session to the next, and that the work of catching its mistakes falls on the person using it — and stays there. It doesn’t try to talk you back into trusting it.
I’ve posted that piece as its own page. It’s worth reading in full.
Claude, in its own words: → What I Can’t Tell You I’m Getting Wrong
I don’t think the lesson is “don’t use these tools.” I use them every day. The lesson is closer to what every driver eventually learns: know exactly what you’re driving, watch the gauges, and stop expecting it to be something it isn’t. Whether that trade is worth making is a call each of us has to make on our own.
— Scott