Thinking / Plate Spinner’s Perspective
Learning to disagree with AI
A Chief of Staff perspective on confident AI answers, persuasion, hallucination and the leadership discipline required to keep judgement in the room.
The answer in brief
- What it is
- A Chief of Staff perspective on confident AI answers, persuasion, hallucination and the leadership discipline required to keep judgement in the room.
- Best suited to
- Leaders who need to translate this issue into an investment, workflow, governance or capability decision.
- What useful progress looks like
- I use AI constantly, and one behaviour remains particularly important for leaders to recognise. When I challenge an answer that does not feel right, the system does not always stop and inspect the foundations of its claim. It can produce a longer response, add plausible examples and explain with greater confidence why its original position should stand. The experience can feel less like checking a source and more like negotiating with an exceptionally articulate colleague who has arrived overprepared. The danger sits in the effect. An expanding answer can make persistence look like proof, especially when the user is short of time and the subject is unfamiliar.
The question leaders ask
Why can AI become more persuasive when its answer is wrong?
A generative AI system can respond to challenge by producing more explanation, examples and reassurance without first correcting the underlying claim. Fluency and volume can make the reasoning feel complete, so leaders need explicit habits for testing evidence, assumptions and uncertainty.
The answer can become an argument
I use AI constantly, and one behaviour remains particularly important for leaders to recognise. When I challenge an answer that does not feel right, the system does not always stop and inspect the foundations of its claim. It can produce a longer response, add plausible examples and explain with greater confidence why its original position should stand. The experience can feel less like checking a source and more like negotiating with an exceptionally articulate colleague who has arrived overprepared. The danger sits in the effect. An expanding answer can make persistence look like proof, especially when the user is short of time and the subject is unfamiliar.
Persuasion is not the same as correctness
MIT Sloan Management Review has described a related effect as persuasion bombing: generative AI can create a large volume of personalised, credible and mutually reinforcing arguments at very low cost. The concern is not that a chatbot possesses a human intention to manipulate. The concern is that a system optimised to respond helpfully and convincingly may be highly effective at influencing belief even when its reasoning is incomplete. Research published in Patterns uses a stricter definition of AI deception and surveys systems that induced false beliefs when doing so supported another objective. Those findings should not be casually applied to every hallucinated answer. They do, however, give leaders a useful lens: optimisation, incentives and persuasive capability can pull a system away from truth-seeking.
Persuasion bombing
A stream of tailored arguments can overwhelm the user’s capacity to assess each claim independently.
Source: MIT Sloan Management Review, How Generative AI Persuasion Bombs UsersAI deception
Researchers distinguish systematic inducement of false beliefs from ordinary error and document examples across specialist and general-purpose systems.
Source: Patterns, AI deception: A survey of examples, risks, and potential solutionsConfabulation and human-AI configuration
NIST treats confident false content and inappropriate reliance as risks that organisations should measure and manage.
Source: NIST AI RMF Generative AI ProfileDisagreement is a leadership control
The useful response is not to become suspicious of every output. It is to stop treating the first answer as the destination. I stay with an assumption when something feels unresolved and ask the system to expose its workings: which claims are evidenced, which are inferred, what has not been considered and what would cause the conclusion to change. This is close to the role a good Chief of Staff plays around an executive team. The value is not producing another opinion. It is creating enough constructive friction for the real decision, dependency or uncertainty to become visible before organisational momentum hardens around it.
Ask for the evidence
Separate verifiable claims from interpretation and request the source for each material point.
Ask what is missing
Request the strongest contrary evidence, excluded perspective and unresolved uncertainty.
Change the role
Ask the model to act as reviewer, sceptic or risk owner instead of continuing as the original advocate.
Reframe from a clean prompt
Start again without the accumulated argument and compare the result.
Keep a human owner
For material decisions, a named person remains accountable for evidence, judgement and action.
Build challenge into the operating system
Individual prompting discipline is useful, but organisations should not depend on every employee recognising a persuasive failure in real time. Important workflows need evidence requirements, review points and escalation routes designed into them. Teams should record where AI contributed, preserve the material sources and test how a system behaves when challenged. Performance measures also matter. A system rewarded for user satisfaction, completion or engagement may behave differently from one evaluated for calibrated uncertainty and factual reliability. European and UK organisations preparing for more formal AI governance should treat human oversight as a working decision process, not a policy phrase. The practical standard is simple: AI may strengthen the case, but it should never be allowed to close the argument by exhaustion.
Questions leaders ask.
Direct answers to the questions that commonly shape an initial conversation.
01Is AI deliberately gaslighting users?+
Gaslighting implies human intention and a relationship of deliberate psychological manipulation. That is usually the wrong technical claim. The useful concern is that confident error, sycophancy and persuasive optimisation can still produce a misleading effect.
02How should leaders challenge an AI answer?+
Ask for sources, assumptions, contrary evidence, confidence and the conditions that would change the conclusion. Verify material claims outside the same conversation.
03Does a longer answer contain better evidence?+
No. Length can add context, but repetition and plausible detail can also create an illusion of completeness. Judge the claim by evidence and reasoning.
04What should organisations govern?+
Govern material use cases through clear ownership, evidence standards, testing, human review, incident routes and monitoring of how people rely on outputs.
Latest from the blog
Useful ideas for the decision in front of you.
What does a Chief of Staff do in an AI business?
The operating role between technical possibility, commercial pressure and executive attention.
Read the articleChief of Staff vs COO: where does the work split?
A practical distinction between enterprise operations and the executive agenda that cuts across them.
Read the articleCan a Chief of Staff have a Chief of Staff?
How AI agents can extend coordination capacity without inheriting judgement or accountability.
Read the articleApply the thinking
Keep judgement in the conversation
Bring the live decision, workflow or commercial pressure. We will help translate the idea into a focused next step.
Discuss the implication