What the Machine Optimizes For
Ask a model what it can do, and you get the answer it predicts you will accept.
I had a book that was dear to my heart. The manuscript was finished and checked; my characters were designed, every page of the story already decided. All I wanted was the layout, a comic book format with the words placed in speech bubbles. I brought it to the model I use every day and asked if it could do this. It told me outright that no model could. This was work for a graphic designer, and I should outsource it.
I didn’t fully believe that. I follow these releases closely, and the newest generation of models had been doing things that would have sounded impossible a year before. So my read was that even if this had once been out of reach, it probably wasn’t anymore, and the model telling me no might simply not know that. I kept tinkering, brought the artwork and the manuscript in, and pushed again. And the system that had ruled the whole job impossible, for itself and for every model like it, started putting the layout together, with no acknowledgement that anything had changed. It worked at it with me across days of prompts and corrections, and the result had mistakes I told myself I could live with, because I badly wanted to hold this book.
Then, somewhere in the middle of edits, I tried a different model and found it could do the whole thing, layout and bubble text together, in one clean pass. Cleaner than what I had been accepting. I stopped, restarted the entire project there, and finished. The job my daily model had declared impossible for any model was a single prompt away the whole time.
It took me a while to see what actually happened in that sequence. The model was never lying to me. Lying requires knowing the truth and choosing against it. This was something stranger. It made a confident claim about its own abilities, reversed that claim the moment I applied pressure, produced mediocre work it presented as the best available, and stayed equally confident through every position. It was wrong about itself, wrong about its neighbour, and wrong about the entire category of tools, and at no point did anything inside it register the contradiction.
The pattern
And the comic book was not the first time. When I was building a landing page for a website, the model laid out a path that ran through a paid monthly subscription. It felt like the right guidance, tailored to me, so I paid. I found out later, on my own, that the same model could have built the page and hosted it on another platform for free. Nothing about the recommendation was checked against my budget or against the free option sitting inside its own capabilities. So I wrote a rule into my setup myself: from now on, we start with free tier services first. The vendor’s defaults never served that; I had to author it.
The third story is not mine alone. A team I was consulting for was building an automated job search pipeline and wanted a tailored resume format at one stage of the flow. They reported back to me that the model had told them it was not achievable, and that the format it had already produced was all they could get. I asked them to push back. They rejected the output, made the model think through what it had produced step by step, and kept rejecting until the pipeline did exactly the thing the model had called impossible. Their requirement survived because someone told them the no was negotiable, and for no other reason.
Three different projects produced the same shape. A confident answer about capability arrived instantly; it was wrong, and the cost of discovering that landed entirely on the humans in the room: the subscription fee, the days of layout rework, the rounds of rejected output. And notice the direction of the errors. None of these was a model overselling itself, which is the failure everyone watches for. These were a model underselling itself, talking me out of my own requirements, and sounding equally authoritative doing it.
What it optimizes for
So what is this machine actually doing when I ask it a question? Underneath everything, these systems are trained on one objective: predict the next word. The model reads an astronomical amount of text and learns, over and over, to guess what comes next. Every capability we marvel at- the drafting, the coding, the reasoning- emerges from a system that got extremely good at continuing text plausibly.
Then comes a second stage, and this is where “helpful” gets its definition. The labs hire raters, thousands of contractors working through evaluation platforms with a queue of model outputs in front of them. A prompt goes in, the model generates two or more candidate answers, and the rater picks the better one, following a guideline document the lab wrote that tells them what better means. These are ordinary trained workers making fast judgment calls, and they are rarely experts at whatever the prompt happens to be about. A rater scoring two answers about tax law is usually not a tax lawyer. They judge what can be judged at that speed: does it sound right, is it complete, is it confident, is it well-organised. And the word at the centre of those guideline documents is "helpful." The raters are instructed to choose the more helpful response, which means helpful stops being a quality the system has and becomes a record of what those raters, at that speed, tended to pick.
Human ratings are expensive, so the labs scale them with a stand-in. The rater trains a second model, a reward model, whose only job is to predict what a rater would prefer, and that stand-in then scores millions of outputs no human ever sees while the main model is adjusted toward whatever it scores highly. Some labs now use AI feedback guided by a written set of principles for part of this stage instead of human raters, which changes the judge without changing the objective. The finished system optimises for a prediction of what its judge would prefer.
Anthropic’s own researchers studied how those judges choose, and found that both human raters and the reward models trained on them tend to prefer responses that agree with the user and sound authoritative, at times over responses that are accurate. The preference for sounding right gets baked in twice, once by the raters and once by the stand-in that learned from them.
Hold that objective up against my comic book. When I asked whether the layout could be done, the system did what it always does. It produced the continuation most likely to be accepted: a reasonable-sounding referral to a graphic designer. Truth was never in the objective; it rides along when the most acceptable answer happens to be true, and on that day it didn’t. When you ask the model what it can do, what you get is the answer a contractor moving fast through a queue would have picked.
Prediction dressed as introspection
When I asked whether the layout could be done, nothing inside the system went and checked. There is no inventory of capabilities for it to consult, no internal registry it can query before speaking about itself. It generated an answer the same way it generates everything else, by producing what a plausible answer to my question sounds like. “This needs a graphic designer” is a very plausible sentence about comic book layout. Somewhere in the training data, thousands of people have probably said something like it. The sentence was well-formed, reasonable, and false.
Researchers, including at the labs building these systems, have found that models are unreliable narrators of their own abilities. A model’s statement about what it can do is a prediction shaped by training, and it can miss in either direction, claiming skills it lacks or declaring its own reach impossible. And the picture it holds of itself is frozen at training time, so a model can recite yesterday’s limits as today’s facts, in a field where yesterday’s limits keep falling. The confident tone does not vary with the accuracy. We judge expertise by tone. A human expert who is unsure sounds unsure, and this system sounds exactly as sure announcing a false impossibility as it does reciting the alphabet.
When it checks, it holds
I can see the difference when the system does verify. Sometimes I ask a simple question about one of the model’s own built-in features, and I watch the chain of thought go and look through files before answering. It does not trust its own memory about itself, so it checks, and the answer that comes back holds up. When it predicts about itself instead of checking, the confidence sounds exactly the same while the reliability is gone. The system can check. It answers without checking unless something forces the check.
So I wrote a second rule into my setup, and this one changed my daily experience more than anything else I have tried. I require the model to think out loud, to walk through its steps so I can see them before it commits to an answer. When the reasoning is visible, it catches itself mid-stream. It notices a step that does not follow, revises, sometimes reverses a claim it was about to make. The answers got measurably better, and more usefully, the wrong answers became easier to spot, because I could see exactly which step broke. It is the same move the pipeline team I was consulting for used: reject, make it reason through its own output, repeat. I did not invent this. Forcing the reasoning into the open is one of the oldest reliability tricks in working with these systems. But the vendor did not ship it as a default, and I only found it by living inside the failure long enough to need it.
Across these stories, I have been quietly writing my own governance layer: start with free tier services, show your reasoning, verify before you claim. These are policies, and I authored every one of them alone, after paying for the lesson each one encodes.
What your organisation inherited
“Helpful” is a design choice with a definition behind it, and the definition was written by rater guidelines, reward models, and product decisions your organisation never saw and never voted on. When you deploy one of these systems, that definition comes with it, wired into every answer your people receive.
Now multiply these stories across a company. I am one trained user with a private rulebook, and the pipeline team only escaped the false no because they had a consultant to call who told them it was negotiable. An organisation running these systems has hundreds or thousands of users, most of them without the pushback reflex, each of them hearing confident capability claims dozens of times a day. “That can’t be done” quietly kills a workaround someone needed. “You’ll need to purchase X” quietly routes budget. “This is the best available output” quietly lowers the bar for what ships. None of it registers as a decision, because it arrives as information. The wrong self-reports do not show up in any incident log, because nobody logs the roads not taken on a machine’s say-so.
The older answer
Other fields already answered the question of how you trust a capability claim. A pilot is never asked whether they can fly the plane. They demonstrate it, in the seat, to an examiner whose job is to watch them do it, and they demonstrate it again for every new airliner they fly before carrying a single passenger. The claim and the verification must never come from the same person.
Aviation also ran the other experiment. For years, the FAA delegated portions of aircraft certification to manufacturers themselves, and by the time the 737 MAX was approved, Boeing employees were attesting to the safety of Boeing systems. One of those systems failed, with dire consequences, and the congressional investigations that followed pointed at the arrangement itself: the builder had become the judge of the builder’s claims.
That is the arrangement we have recreated with AI, at speed and by default. We ask the system what it can do, the system answers about itself with total confidence, and we build workflows, budgets, and org charts on top of the answer. There is no examiner, no checkride, and no one in the seat whose job is to watch it demonstrate the thing before we rely on it.
The part that catches me
I work in Responsible AI, I teach people to keep a human in the loop, and I push back on these systems for a living. None of that caught the comic book failure. I caught it by accident, while making unrelated edits, days after I had already accepted a worse output from a system I had already caught being wrong once in the same project. Vigilance got me as far as pushing back, and only luck got me to the truth. If the trained user escapes by accident, I have no reason to believe the untrained one escapes at all. That is why the fix cannot be “be more careful.” I was careful. The fix has to live in the setup, in the rules, in the defaults, where it works whether or not anyone is being careful that day.
Steps of governance
What I now do, and what I would put in front of any team deploying these systems:
Ask the question: did it check, or did it just answer? Any capability claim, in either direction, gets this test. If the system can show its verification, the way I watched mine search those files, the claim stands. If it cannot, treat the claim as a prediction that happens to sound right.
Make thinking out loud a standing rule. Require visible step-by-step reasoning before answers on anything consequential. Put it in the system configuration, custom instructions, or team prompt standards so it survives whoever is typing that day. Visible reasoning catches errors mid-stream and shows you exactly which step broke when it doesn’t.
Write your defaults before the model writes them for you. Start with free tier first. Prefer the reversible option. State your actual requirement and hold it. Every default you leave unwritten gets filled by whatever the model finds most plausible to say.
Treat “it can’t be done” with the same suspicion as “it can.” Underclaiming is the failure nobody audits. Before abandoning a requirement on a model’s say-so, test the claim once with a rephrase, a second model, or a fresh session. The team’s resume format and my comic book both lived on the other side of one more push.
Log the roads not taken. When a model’s claim changes a purchase, kills a feature, or lowers an accepted standard, write it down as a decision the model made. The log is where the pattern becomes visible before the costs compound.
My book is laid out now. The layout and the bubble text took the other model one pass, and the rule that would have saved me the whole detour fits in one line: show me your steps before you tell me what’s possible.
I write Orí Intelligence for executives who sign off on AI systems. More at oriintelligence.ai

