9 min read

Why AI doesn't give you what you want

Almost everyone who works with a language model daily has lived the same scene. You ask for something, you get back a flawless piece of writing, well structured, correct, and useless to you. The instinctive reaction is always the same: turn on reasoning mode, tell it to think harder, consider switching models.

And it almost never works, because the failure was rarely where you were looking.

Behind any request you make to an AI there are three distinct operations, and we treat them as if they were one. Specifying is getting it to understand what you already want. Exploring is having it show you options you didn't know existed. Executing is having it done well.

They are independent: you can nail one and fail badly at another. You can have a perfectly defined request and a disastrous execution. You can get a flawless execution of a request that was never really yours. And you can get both right and never find out that a third, better path existed.

Each one fails differently and each is fixed by a different remedy. The expensive mistake, the one made every single day, is applying one axis's remedy to another axis's failure.

Axis 1: specifying, getting it to understand what you already want

The symptom is instantly recognisable: you ask for something and receive an answer that is correct but generic, the kind that is well written, never commits to anything, and helps you not at all.

It isn't that the model is dim. It's that your request admits many reasonable interpretations, and the model doesn't pick one: it hands you something close to the average of them all. That average is exactly what we recognise as the textbook answer.

Think of a request as ordinary as "put together a training plan for my team". Hundreds of different briefs fit inside that sentence, all of them defensible. Who is it for? How long is it? What technical level does it assume? What format? What tone? How far does the scope go? How will we know it turned out well? The model guesses none of that: it blends everything.

And the failure isn't proportional, it's multiplicative. If the model gets four out of five of those decisions right individually, which is generous, the odds of it getting all of them right at once fall to a little over one in six. That's why the experience is so disorienting: the model looks competent in every sentence, because it is, but you only ever see the aggregate result.

The remedy is to have it ask you questions before it starts. This isn't politeness towards the machine: every answer you give wipes out a whole cluster of interpretations at once.

The instruction that converges

Before you execute, ask me the 3 questions whose answers would most change the final result. Only questions where you can't anticipate my answer, that I'm able to answer, and that aren't already covered in this brief. Wait for my answers before you start.

Every line is doing concrete work. "Would most change the result" rules out irrelevant detail. "Where you can't anticipate my answer" stops it asking what it already knows. "That I'm able to answer" prevents that awkward interrogation about things you don't know either. And the limit of three is what separates a useful conversation from a form to fill in.

A few well-aimed questions beat many.

Asking first is not the same as correcting afterwards

The reasonable objection is "I'll just correct it in the next message". Same information, different order, what difference does it make?

A lot. These models generate word by word, conditioned by everything they have already written. The moment they lock in an interpretation in the opening paragraphs, that interpretation becomes part of the context and skews everything that follows. The error doesn't dilute across the text: it compounds.

When you ask first, the information lands in a clean context. When you correct afterwards, your correction is competing against two thousand words of already-written commitment. It isn't the same move.

There is one case where none of this applies: if what you're asking for has a single correct answer, a closed calculation or a literal translation, there is nothing to ask about. There, questions are pure cost.

Axis 2: exploring, having it show you what you didn't know was there

Here is the part almost nobody separates out, and it's the most valuable.

Everything above assumes you know what you want and the problem is conveying it. In a good share of real cases that's false. When the subject is complex and you're just starting, you don't have a clear intention: you have a vague sense of direction. There is nothing to convey, because you haven't formed it yet.

What you need from the model then isn't for it to guess what you want, but to show you the map of the terrain so you can work out where you want to be.

An example that makes it obvious

Someone has to run four internal training sessions on a tool the team has just started using. They brief it with all the detail in place: fourteen non-technical people, ninety-minute sessions, no theory, they should walk out able to do three specific things, delivered as a script for whoever runs it.

It's impeccably specified. With the converging instruction above, the model comes back with three reasonable questions: whether the sessions are consecutive or weekly, whether there's a test environment available, whether the person running them already knows the tool or will be learning it on the fly. All three improve the material. None of them touches the substance.

With an exploration instruction, something very different surfaces. Among the decisions the model flags as critical there is one nobody had formulated: is the team's problem that they don't know how to use the tool, or that they don't know when to use it?

Those are two incompatible courses. One is product training: clicks, features, guided practice. The other is decision criteria: which situations this tool is right for and which ones will cost you an afternoon. In four ninety-minute sessions you can do one of them well, or both badly.

And here's the interesting part: no amount of better specification would have prevented it. The person who wrote the brief didn't leave that out in a rush. They left it out because it hadn't occurred to them that it was a decision. Their brief was correct and complete within the frame they had in their head, and the frame was the thing that was wrong.

A better model wouldn't have saved it either. Given the original brief, any competent model produces a good product course, flawlessly executed. It simply wasn't the course that was needed.

The pattern repeats everywhere. Redesigning a pricing page when the problem was the pricing model. Migrating the documentation when the right move was deleting seventy per cent of it. The brief is correct, the execution would be good, and the real decision was one level up.

The instruction that diverges

Don't execute yet. First: tell me which 4 or 5 decisions this brief actually turns on, even ones I haven't mentioned; for each of them, give me the real alternatives, including one that people in my sector don't usually consider; and flag which ones are mutually incompatible. Then wait for me to choose.

That third part is the one most often forgotten and the one that contributes most. Knowing which options rule each other out tells you where the real decision sits, instead of leaving you picking between fifteen cosmetic variants.

A limit worth knowing about

The model will hand you the conventional alternatives from your domain: the ones a competent professional in your sector would have mentioned. That makes it an excellent inventory of omissions, and that is the part that pays off almost every time.

It works considerably worse as a source of originality. By construction it returns the centre of what has already been thought. If you're expecting the idea nobody had had, you're asking for the exact opposite of what its mechanism produces. It's still worth doing, as long as you don't mistake covering the known space for reaching the new one.

Axis 3: executing, having it done well

If the model is going to invent a figure, get a calculation wrong, or cite a source that doesn't say what it thinks it says, it will do so just the same with a perfect brief and five alternatives on the table.

Asking questions protects you from solving the wrong problem well. It doesn't protect you from solving the right problem badly. Verification is still yours, and there's no shortcut.

Here the remedy genuinely is different: extended reasoning, verification, breaking the work into steps. And it works, when the failure belongs to this axis.

The trouble is that people reach for reasoning when what failed was the specification. And there it doesn't just fail to help: it makes things worse. The model doesn't create the information it's missing; what it does is build a more elaborate and more convincing justification for the wrong assumption. The result is worse than the original, because now it's far harder to spot that it set off from the wrong place.

Two kinds of question that aren't the same thing

A specification question is discriminating: it separates hypotheses, closes an option, converges. "Word or PDF?" is perfect for that.

An exploration question is generative: it opens territory, it surprises you, it diverges. "Which decisions are at stake that I haven't mentioned?" is the other kind.

The optimal question for one axis is a mediocre question for the other. That's why a single instruction along the lines of "ask me three questions" only covers half the job.

A note on the tools that have turned this into option buttons is worth the detour. They're good for specifying and weak for exploring: they discriminate well between alternatives already on the table, but they can't hand you the one that isn't. And they carry a non-obvious risk: if none of the options is what you wanted and you pick the least bad one, you have just handed the model a wrong belief with full confidence attached. That's worse than not asking, because now neither of you will ever revisit that decision.

Diagnosing what went wrong

The useful question isn't "which model should I use" or "how do I prompt it better". It's which axis failed. There are three symptoms, and each has its remedy, along with the remedy that backfires.

  • It's well made, but it isn't what I wanted. A specification failure. Fixed by: discriminating questions before you start. Made worse by: asking it to reason harder.
  • It's what I asked for and it's good, but reading it I can see there was a better path. An exploration failure. Fixed by: generative questions, with the incompatibilities flagged. Made worse by: specifying harder, which would only have you executing sooner in the wrong direction.
  • It was what I wanted, the path was right, and it's badly done. An execution failure. Fixed by: extended reasoning, verification and breaking it into steps. Made worse by: more questions, which would have changed nothing.

The expensive mistake is the crossover: applying the third axis's remedy to a first-axis failure. It happens daily, it's what a good chunk of AI commentary promises, and it produces answers that are longer, more self-assured and exactly as useless.

What's almost never missing is capability

If you take one idea away, make it this one: when something goes wrong, the instinctive reaction is to tell the model to think harder, and that mostly succeeds at making the wrong assumption more convincing.

What's missing is almost never capability. It's a piece of information still sitting in your head, or an alternative nobody has put on the table.

One last note for anyone building autonomous agents: if the agent asks outwards, towards other systems rather than towards a person, that clarification channel widens the attack surface. Someone else can answer on your behalf. In a normal conversation, with a person on the other side, this doesn't apply.

Found it useful? Share it with your team.