Measuring AI value after the impressive demonstration
Part of a retrospective. The retrospective month groups the topic; it is not an earlier publication date. Published 18 September 2026.
Measure completed work, review effort and attributable outcomes so an impressive AI demonstration does not become an unsupported savings claim.
An AI demonstration can show that a task is possible. It cannot, on its own, show that the organisation is better off doing the task that way.
The missing work usually sits around the generated output. Someone gathers the input, repairs the format, checks the answer, corrects omissions and moves the result into the next system. A fast draft can coexist with a slow process.
I would measure value at the point where the business accepts the completed work. That keeps the evaluation attached to the outcome rather than the speed of the most visible model call.
Establish the comparison before the pilot
Choose a recognisable unit of work. It might be a completed brief, a resolved request or a reviewed document extraction. Define what acceptable completion means so that the manual and AI-assisted routes face the same standard.
Observe the existing process. Include preparation, waiting, review, corrections and handover. Distinguish elapsed time from staff effort: a task can spend a long time in a queue without consuming continuous labour.
Record the variation rather than one convenient average. A straightforward request and a difficult exception may require very different effort. If the pilot receives only the simple cases, comparing it with the entire historical workload will exaggerate the benefit.
A hypothetical drafting pilot could compare briefs of similar length, source quality and complexity. If the AI group receives better templates and cleaner information, record those changes too. They may be worthwhile improvements, but the model should not receive all the credit.
Do this before people become invested in the result. Retrospectively choosing a baseline after a successful demonstration invites a comparison that is technically true and commercially misleading.
Count the work that moves elsewhere
Generation time is only one part of the process. Measure input preparation, prompt revision, verification, correction and final acceptance. Include the extra effort created when the assistant returns something plausible but incomplete.
Ask who now does that work. An apparent saving for the original author may transfer effort to a senior reviewer. If the reviewer is a scarce resource, the organisation may have moved its bottleneck rather than reduced it.
Capture rework after acceptance as well. A defect discovered by the next team belongs to the AI-assisted process even if the initial user clicked approve. Use an agreed observation window and make its limits clear.
Keep quality categories visible. A spelling correction is different from an omitted contractual condition or a disclosure to the wrong recipient. A single error count can hide a worse mix of consequences.
The purpose is not to make the pilot fail by counting every possible cost. It is to compare both routes honestly. Existing manual processes also have errors, review burdens and workarounds; those belong in the baseline.
Separate capacity from cash
Time released is not automatically money saved. Staff may use it to reduce a backlog, improve quality or take on additional work. Those can be real benefits without a corresponding reduction in payroll.
I would report released capacity separately from cash effects. A cash-saving claim needs an identifiable change in expenditure, such as an avoided external cost or a budget adjustment the owner is willing to make. Do not multiply estimated minutes by salary and present the result as money already recovered.
Likewise, a claim about faster decisions needs evidence that the decision happened sooner, not merely that the briefing draft arrived earlier. Approval schedules or unresolved information may still determine the timeline.
Where benefits are estimates, label the assumptions. Show the workload to which they apply, expected usage, quality constraints and the operational change needed to realise them. Avoid a precise-looking annual figure built on a short, unusually enthusiastic pilot.
For uncertain benefits, a range of scenarios may be more useful than a single forecast. Those scenarios are planning assumptions until observed results replace them.
Treat adoption as a diagnostic, not the outcome
Usage helps explain whether a tool fits the work. It does not establish value by itself. High prompt volume may reflect genuine usefulness, repeated failed attempts or a task that now needs more interaction than before.
Look at completed tasks and repeat use by the intended users. Ask why others stopped. The answer may be poor source coverage, slow responses, unclear data rules or an interface that interrupts the workflow.
Avoid coercing usage to improve the pilot's statistics. If staff must use the tool, record that condition. Mandated activity should not be presented as voluntary demand.
Interview users about specific recent tasks rather than asking whether they “like AI”. Compare their account with the process evidence. A person may enjoy the assistant while spending longer checking it; another may dislike the interface yet produce measurably better work.
Privacy matters here too. Collect only the information needed to understand performance, explain what is being measured and avoid turning a value study into undisclosed individual surveillance.
Make the funding decision explicit
The pilot report should separate observed results, estimated benefits and unresolved questions. It should also include recurring operating costs: support, evaluation, source maintenance, service usage and change work. The hidden operating costs of an internal AI platform are part of the value calculation, not a later infrastructure discussion.
Agree the decision rules with the sponsor. A useful outcome may justify expansion; a quality problem may require a narrower scope; insufficient evidence may justify another bounded trial. An improvement that depends on unsustainable expert review may not justify production.
Keep the comparison alive after launch. Usage patterns change, source material grows and service changes can alter quality or effort. A benefit demonstrated in a controlled pilot is a claim to maintain, not a permanent property of the tool.
I would rather fund a modest, evidenced improvement than a dramatic savings estimate nobody can reconcile to completed work. That makes the next investment discussion easier, because the organisation knows what it is buying and how it will notice if the benefit disappears.
Working on something similar?
If this connects with something you’re working through, I’m happy to talk about where you’re stuck and whether I can help.
See how I work