Thinking / Data readiness
What good enough data looks like for an AI pilot
How to set a proportionate information standard that supports meaningful evidence without waiting for perfect enterprise data.
The answer in brief
- What it is
- How to set a proportionate information standard that supports meaningful evidence without waiting for perfect enterprise data.
- Best suited to
- Leaders who need to translate this issue into an investment, workflow, governance or capability decision.
- What useful progress looks like
- Leaders can postpone useful experimentation while waiting for a broad data transformation to finish. Others run a demonstration on a carefully curated sample that bears little resemblance to production. Both approaches weaken the investment decision. A good pilot uses information close enough to real operating conditions to reveal quality, access and governance constraints. It narrows the workflow and consequence so those constraints remain safe and manageable.
The question leaders ask
Does data need to be perfect before an AI pilot?
No. It must be representative, permitted, sufficiently accurate and traceable for the hypothesis being tested. Known limitations should be documented and reflected in the pilot's scope and human controls.
Perfection delays learning
Leaders can postpone useful experimentation while waiting for a broad data transformation to finish. Others run a demonstration on a carefully curated sample that bears little resemblance to production. Both approaches weaken the investment decision. A good pilot uses information close enough to real operating conditions to reveal quality, access and governance constraints. It narrows the workflow and consequence so those constraints remain safe and manageable.
Define the standard against the hypothesis
The required data conditions should follow what the pilot needs to prove. A search assistant may test retrieval accuracy and user trust. A recommendation product may require historical outcomes, consistent entities and bias assessment. A workflow agent needs reliable state, permissions and action logs.
Representative
The sample reflects the variety, ambiguity and exceptions present in normal work.
Current
The system can identify the active version and prevent superseded material dominating results.
Permitted
Access, confidentiality, privacy and intellectual-property conditions are explicit.
Traceable
Outputs can be linked to the sources, transformations and versions that influenced them.
Measurable
A test set and baseline allow the team to judge quality rather than admire fluency.
Recoverable
Errors can be detected, contained, corrected and learned from.
Make limitations visible
Document missing sources, known error patterns, excluded users and assumptions about future integration. Tell pilot users what the system can and cannot support. Record when they correct, override or abandon an answer. This evidence prevents a controlled trial from being interpreted as proof of universal readiness. It also helps the team estimate the information, engineering and change work required for production.
Convert discovery into a data backlog
Each failure should become a specific requirement with an owner and relationship to business value. Examples include retiring an outdated policy, reconciling customer identifiers, improving document metadata or creating a controlled knowledge source. Prioritise the fixes that materially change performance in the chosen workflow. This creates a defensible foundation roadmap built from product evidence rather than a generic ambition to clean everything.
Questions leaders ask.
Direct answers to the questions that commonly shape an initial conversation.
01How much data does a pilot need?+
Enough representative material to test the important variations and failure modes. The quantity depends on the product, workflow and required confidence.
02Can synthetic data be used?+
It can support early testing where privacy or availability is constrained. Production confidence still requires representative real-world conditions and careful validation.
03Who approves pilot data?+
The business owner, data owner and relevant privacy, security or risk specialists should agree purpose, access, retention and controls.
04What if the pilot reveals poor data quality?+
Treat that as valuable evidence. Narrow, pause or redesign the product, then prioritise the specific foundation work required.
Latest from the blog
Useful ideas for the decision in front of you.
Learning to disagree with AI
Why a confident, persuasive answer should be the start of leadership judgement, not the end of it.
Read the articleWhat does a Chief of Staff do in an AI business?
The operating role between technical possibility, commercial pressure and executive attention.
Read the articleChief of Staff vs COO: where does the work split?
A practical distinction between enterprise operations and the executive agenda that cuts across them.
Read the articleApply the thinking
Define a proof that reveals what must be true
Bring the live decision, workflow or commercial pressure. We will help translate the idea into a focused next step.
Discuss the implication