Thinking / Data readiness

What good enough data looks like for an AI pilot

How to set a proportionate information standard that supports meaningful evidence without waiting for perfect enterprise data.

The answer in brief

What it is
How to set a proportionate information standard that supports meaningful evidence without waiting for perfect enterprise data.
Best suited to
Leaders who need to translate this issue into an investment, workflow, governance or capability decision.
What useful progress looks like
Leaders can postpone useful experimentation while waiting for a broad data transformation to finish. Others run a demonstration on a carefully curated sample that bears little resemblance to production. Both approaches weaken the investment decision. A good pilot uses information close enough to real operating conditions to reveal quality, access and governance constraints. It narrows the workflow and consequence so those constraints remain safe and manageable.

The question leaders ask

Does data need to be perfect before an AI pilot?

No. It must be representative, permitted, sufficiently accurate and traceable for the hypothesis being tested. Known limitations should be documented and reflected in the pilot's scope and human controls.

ScopeRepresentative subset
StandardKnown and controlled
OutputEvidence plus data backlog
01

Perfection delays learning

Leaders can postpone useful experimentation while waiting for a broad data transformation to finish. Others run a demonstration on a carefully curated sample that bears little resemblance to production. Both approaches weaken the investment decision. A good pilot uses information close enough to real operating conditions to reveal quality, access and governance constraints. It narrows the workflow and consequence so those constraints remain safe and manageable.

02

Define the standard against the hypothesis

The required data conditions should follow what the pilot needs to prove. A search assistant may test retrieval accuracy and user trust. A recommendation product may require historical outcomes, consistent entities and bias assessment. A workflow agent needs reliable state, permissions and action logs.

01

Representative

The sample reflects the variety, ambiguity and exceptions present in normal work.

02

Current

The system can identify the active version and prevent superseded material dominating results.

03

Permitted

Access, confidentiality, privacy and intellectual-property conditions are explicit.

04

Traceable

Outputs can be linked to the sources, transformations and versions that influenced them.

05

Measurable

A test set and baseline allow the team to judge quality rather than admire fluency.

06

Recoverable

Errors can be detected, contained, corrected and learned from.

03

Make limitations visible

Document missing sources, known error patterns, excluded users and assumptions about future integration. Tell pilot users what the system can and cannot support. Record when they correct, override or abandon an answer. This evidence prevents a controlled trial from being interpreted as proof of universal readiness. It also helps the team estimate the information, engineering and change work required for production.

04

Convert discovery into a data backlog

Each failure should become a specific requirement with an owner and relationship to business value. Examples include retiring an outdated policy, reconciling customer identifiers, improving document metadata or creating a controlled knowledge source. Prioritise the fixes that materially change performance in the chosen workflow. This creates a defensible foundation roadmap built from product evidence rather than a generic ambition to clean everything.

FAQ

Questions leaders ask.

Direct answers to the questions that commonly shape an initial conversation.

01How much data does a pilot need?+

Enough representative material to test the important variations and failure modes. The quantity depends on the product, workflow and required confidence.

02Can synthetic data be used?+

It can support early testing where privacy or availability is constrained. Production confidence still requires representative real-world conditions and careful validation.

03Who approves pilot data?+

The business owner, data owner and relevant privacy, security or risk specialists should agree purpose, access, retention and controls.

04What if the pilot reveals poor data quality?+

Treat that as valuable evidence. Narrow, pause or redesign the product, then prioritise the specific foundation work required.

Latest from the blog

Useful ideas for the decision in front of you.

View the blog

Apply the thinking

Define a proof that reveals what must be true

Bring the live decision, workflow or commercial pressure. We will help translate the idea into a focused next step.

Discuss the implication