← All writing
PublishedRetrospective · July 20265 min readDeliveryLeadership

Evaluating an AI proposal without being distracted by the demo

Part of a retrospective. The retrospective month groups the topic; it is not an earlier publication date. Published 18 September 2026.

Evaluate an AI proposal through the work, evidence, permissions and ongoing cost it would introduce, rather than its best demonstration.

The demonstration answers a difficult question in seconds. The room is impressed, and discussion moves quickly to licences and rollout. Yet the demonstration may have said little about whether the system can perform the organisation's task with its actual information and access boundaries.

I would keep the demonstration, but change what follows it. Ask the supplier to explain the proposed workflow, the evidence behind its claims and the work required when the answer is wrong, unavailable or too uncertain to use.

Start with the decision or task

Define what the organisation wants a person to accomplish. "An AI assistant for staff" is too broad to evaluate. A more testable proposal might help a coordinator locate the current procedure and identify the section relevant to a routine request.

That fictional use case has a clear boundary. Locating information differs from interpreting policy, recommending an exception or taking action in another system. Each added capability changes the consequences and the evidence needed before use.

Compare the proposal with simpler options. Better search, cleaner source documents or a redesigned form may address much of the problem. An AI system may still be useful, but it should earn its place against those alternatives rather than being evaluated in isolation.

Identify the business owner who will judge the result. Technical staff can assess integration and operation, but someone who understands the task must decide what a usable answer looks like and which errors matter most.

Replace curated examples with an evaluation set

Ask to test representative cases selected by the organisation. Include ordinary questions, ambiguous requests, missing information and material conflicts between sources. Keep part of the set separate from the supplier's tuning work so the review is not limited to rehearsed examples.

Define the expected behaviour for each class of case. Sometimes the correct response is to decline to answer, ask for clarification or direct the user to an authorised person. A system that always produces an answer may be unsuitable for a task where uncertainty should interrupt the process.

Evaluate the evidence presented with an answer. A link should lead to material that supports the claim, not merely a document on the same subject. Check versions and permissions. The test should examine whether the user can verify the answer without spending more effort than the existing process required.

Avoid relying on one aggregate accuracy score. Errors have different consequences, and a high average can hide poor performance on a small but important category. Describe coverage and limitations, especially when the evaluation set is small.

Retest after material changes to the model, prompts, retrieval or source collection. Evidence about one configuration should not be carried forward as though the service never changes.

Inspect access and information handling

Map what information enters the service, where it is processed and what is retained. Review supplier terms and the actual configuration with the relevant specialists. A general assurance about privacy is not a substitute for understanding the proposed data flow.

Test with accounts that have different permissions. The system should not disclose restricted material through an answer, citation or generated summary. Examine how permission changes reach the service and what happens while those changes propagate.

Keep the evaluation environment appropriate for the information used. Begin with prepared or approved data where possible. Do not provide a confidential corpus simply because a supplier says it needs realistic material for a demonstration.

If the proposal can take actions, assess those separately from answering questions. Sending a message, changing a record or initiating a payment creates different failure modes. Limit authority and make consequential approvals explicit. Human review needs evidence, time and an enforceable boundary to be useful.

Cost the service around the model

Model usage is only one part of the operating cost. Include source preparation, access integration, evaluation, monitoring, support and specialist review of difficult outputs. Ask who maintains the content and who investigates reports of misleading answers.

A low demonstration cost may not represent normal use. Longer inputs, repeated attempts and additional checking can change demand. Use bounded scenarios and label unverified estimates. Avoid promising a productivity multiplier before observing the work, including correction and review.

Clarify supplier dependencies. What happens when a model is withdrawn, pricing changes or an external service is unavailable? Determine whether the organisation can retain its source material, evaluation cases and configuration in a usable form if it changes provider.

An operating owner should be able to pause a capability when evidence shows it is causing problems. Agree the communication and fallback process so staff know how to continue their work.

Make the next commitment smaller than the uncertainty

A promising evaluation can justify a bounded pilot rather than unrestricted rollout. State the permitted tasks, participants, data and exit decision. Keep unresolved questions visible instead of treating enthusiasm as approval to expand the scope.

NIST's AI Risk Management Framework is voluntary guidance for considering trustworthiness across design, use and evaluation. It can help structure the questions, but citing a framework does not establish that a particular product is suitable or compliant.

The CIO's role remains broader than selecting the model. It includes the process being changed, the people expected to use it and the service the organisation must operate. That is the perspective behind using AI as an IT leader.

A strong proposal should survive discussion of limitations. It should leave the organisation knowing what it is prepared to test, what remains outside scope and what evidence would justify the next investment. The demonstration is useful when it opens that conversation rather than ending it.

Working on something similar?

If this connects with something you’re working through, I’m happy to talk about where you’re stuck and whether I can help.

See how I work