At a glance
- Interruption rules, evidence sufficiency and progression criteria answer different questions.
- Thresholds need a rationale and an agreed approach to uncertainty.
- A dated decision record makes amendments and the next bounded commitment traceable.
An AI pilot needs a decision that remains possible when its results disappoint. Before starting, I would therefore record which observation would justify expansion, which would require further investigation and which would interrupt the pilot. This agreement protects the participants’ limited time. It also makes clear how much uncertainty the next commitment can accommodate.
The decision record below is intended for a bounded organisational pilot. It adds a different question to the detailed evaluation of quality and effort: what decision should follow from the observations? The example is entirely fictional. It concerns preparation of internal procurement material, with no patient care or assessment for medical regulatory approval.
Name the next decision before seeing the first result
A hospital group wants to use AI to turn approved technical documents into comparable summaries for procurement. The hoped-for improvement is a usable decision brief produced with less total effort. After four weeks, the group will decide whether to test the same process in another procurement team. A group-wide rollout is explicitly outside that decision. The duration and scope are chosen assumptions for this example.
This boundary changes the justification required. Another small test needs different evidence from a long-term contract or automated external communication. The record should therefore identify the next step, its additional costs and the person authorised to decide. Saying “we will see what comes next” leaves these questions open and commits nobody to a decision that the pilot can inform.
Separate three kinds of rule
First, the pilot needs interruption rules for events that breach its permitted scope. In this example, these would include sending material to an unapproved destination or using an unreviewed draft as a binding purchase order. The interruption applies to the agreed use. The responsible person investigates the event and decides on further action through the existing procedure.
Second, there are rules about sufficient evidence. If the pilot handles only unusually simple documents, it may omit the very cases that justified testing the application. Missing time records can also make the comparison unusable. Such a finding initially means that the planned decision lacks a sufficient basis. It establishes neither the application’s benefit nor its unsuitability.
Third, there are rules for economic and organisational progression. These might cover acceptable total effort, adequate quality for the intended purpose and operational support that can actually function. They apply once the permitted boundaries have been respected and the comparison is sufficiently informative. In this structure, a favourable time figure cannot compensate for an unresolved breach of the pilot’s rules.
The 2016 CONSORT extension covers prespecified progression criteria for randomised pilot and feasibility trials. It also cautions against rigid thresholds when estimates are uncertain. Its scope is preparatory research; my organisational decision record is a separate adaptation. [1]
Connect the intended benefit to possible side effects
IHI distinguishes outcome, process and balancing measures. Balancing measures ask whether an improvement creates problems elsewhere. This guidance does not establish the effectiveness of the AI process proposed here. [2]
For the fictional procurement pilot, I would use active working time until a summary is accepted by the qualified reviewer as the outcome measure. This includes preparation, checking, questions and rework across every involved function. A process measure shows whether the agreed review actually happened. A balancing measure captures additional questions reaching specialist departments after the handover.
For every measure, record who collects it, the period it covers and how missing values will be handled. An empty field must not subsequently count as an error-free case or a time saving. Also record which cases were excluded. Results from complete, standard documents explain little about later use with conflicting or incomplete information.
Derive thresholds from the decision
Suppose the team sets a target of at least two minutes less work, on average, for each fully reviewed case before another test is justified. This figure is a freely chosen assumption. Its rationale could be that a smaller reduction would not justify additional support effort in this narrowly defined process. Another team could reasonably choose a different value.
The target also needs a way of handling uncertainty. The team might agree in advance that an apparent improvement with highly variable results triggers a bounded additional measurement period. Clear deterioration in sufficiently comparable cases would argue against expansion. A suitably qualified person should determine the range of uncertainty relevant to the decision and how to estimate it before interpreting results.
Avoid choosing an arbitrary minimum number of cases that suggests scientific certainty. The number required depends on variation, rare error types, differences between reviewers and the consequences of the next decision. Repeatedly processing the same document does not provide independent experience with new cases. If these questions remain unresolved, the record should describe the pilot’s limits and the additional assessment needed.
Complete the decision record together
The printable companion condenses this proposal into nine fields. It supports discussion and documentation. Required professional, technical and legal approvals remain within their established procedures.
Download the one-page decision record as a PDF
- Decision: what bounded next step is being considered, who decides and on what date?
- Task: where does the process begin and end, which cases are included and which are excluded?
- Comparator: which existing process provides the reference, and how will cases be made comparable?
- Expected benefit: what change matters to the receiving function, how will it be measured and why is the target meaningful?
- Quality and side effects: which defects are recorded separately, and what additional work might arise elsewhere?
- Interruption: which event immediately stops the current test, who can stop it and which safe fallback applies?
- Evidence: which cases, measurements and records must be available, and when is the evaluation insufficient?
- Decision rule: what supports expansion, a bounded further check or ending the pilot, and who can justify a departure from the rule?
- Execution: which version is being tested, what time and cost limits apply, where are records held and who follows up?
Rehearse the decision in advance
Before starting, I would test the completed record against three fictional results. In the first, handling time falls but the next department receives more questions. In the second, effort stays the same while the expected quality is achieved. In the third, the few recorded values look promising but difficult documents are missing. Can everyone identify the same intended next step in each case?
Different answers reveal something useful. One person may mean better comparability when they say success; another may expect cost relief alone. This difference should be resolved before the pilot. Otherwise, the team may later argue about data while judging different objectives. The record can expose that conflict, although it cannot force a resolution.
Allow changes that remain traceable
New information may make a previously agreed rule inappropriate. In that case, add a dated amendment: what changed, who agreed, which results were already known and which conclusion consequently becomes weaker? Keep the original version identifiable. An adjustment can be justified while also limiting comparability.
On the decision date, assess observations against the agreed rules. Give unresolved measurement gaps a name and a bounded follow-up task. A positive decision should also specify its scope and the next review point. For me, this is an essential part of entrepreneurial judgement: translating a technical promise into a commitment whose size fits the evidence available.
Sources and further reading
- Eldridge and colleagues: CONSORT extension for randomised pilot trialsPilot and Feasibility Studies
Prespecified progression criteria and uncertainty of estimates in preparatory randomised studies; no evaluation of the organisational record proposed here.
- IHI: outcome, process and balancing measuresInstitute for Healthcare Improvement
Distinction between measurement types and assessment of possible disadvantages elsewhere in a process; no evidence that the proposed AI procedures are effective.
Perspective and interests
This article was developed with AI assistance. All organisational examples are fictional. The practical decision rules are my proposals and have not been tested for effectiveness here.
I am the founder and CEO of aiomics and have a commercial interest in responsible adoption of AI in medicine.



