← Ideas & guides

Reflection

Who is missing from your successful AI pilot?

Early volunteers help a team learn. Wider adoption also depends on who received an invitation, could actually participate and remained visible in the results.

By Dr. Sven JungmannPublished: · Reviewed:
A wooden spoon holds three marbles beside a glass bowl containing many more.

At a glance

  • Selection into a pilot and missing follow-up responses are different questions.
  • Many documents from few users do not replace a broad workforce evaluation.
  • The next group should investigate an important condition that remains unobserved.

A fictional hospital tests an AI documentation tool with a small group of enthusiastic clinicians. They receive protected introduction time and can reach the project team directly. Several report that the tool helps them. Leadership now considers extending access across the hospital. Before interpreting the result, I would ask who never entered the group whose experience is being discussed.

That question protects the value of the pilot. Early volunteers can identify useful applications and expose faults while a service is still being adjusted. Their experience becomes harder to interpret when a decision about everyone else silently rests on it. The next investment should depend partly on which people and working conditions the first test actually covered.

The first willing users answer a bounded question

Ankit Gupta’s Y Combinator talk, published on 14 January 2026, discusses finding a product’s first willing users. It prompted this reflection about the boundary between early participation and wider adoption. The talk offers entrepreneurial advice and contains no clinical evaluation. [1]

For a founder, a colleague willing to persist through initial difficulties is valuable. For a hospital considering routine deployment, that willingness is also part of the test conditions. A clinician who reorganises appointments, practices at home or contacts the supplier personally may make a tool usable in ways the organisation has yet to provide routinely. These are possible mechanisms to investigate. I am not attributing that behaviour to the fictional participants.

I would therefore state two separate questions at the outset: can this group find a worthwhile use, and under what conditions could the intended workforce use it? A positive answer to the first can justify investigating the second. It leaves the second open.

Follow the people into the result

Franklin’s April 2026 report examined 2024 outpatient pilots at two US academic centres. Leadership-selected volunteer clinicians received 166 initial survey invitations; 120 follow-up responses entered analysis. Of respondents, 56.7% called themselves early technology adopters. These surveys describe selected experiences without randomized comparison. [2]

Those figures do not reveal how many employees across the institutions could have benefited. They also do not establish whether absent colleagues would fare better or worse. Selection into a pilot and missing follow-up responses are different questions: who was given an opportunity, and whose experience subsequently became visible? Neither question can be settled by enthusiasm among the people who answered.

For a local decision, I propose keeping a short account of that path. Start with the workforce for whom an expansion is being considered. Record who met the agreed eligibility criteria, who received an invitation, who accepted, who obtained working access and who used the tool for the relevant task. Then show whose outcomes were available at the planned follow-up. Each number should refer to the same defined period or explain the difference.

The denominator changes the meaning of a percentage. A favourable response among people who completed a survey answers a different question from participation among everyone invited. Regular use among people with functioning access differs from regular use among everyone offered a licence. Presenting both can reveal an access problem that an average satisfaction score would conceal.

This organisational proposal has not been evaluated as a measurement instrument. A list of stages also cannot repair a weak comparison or establish clinical benefit. Its purpose is to keep an expansion decision connected to the people actually observed.

Count staff and patients separately

The RE-AIM framework distinguishes participation by intended beneficiaries from adoption by settings and staff. Its official questions, revised in September 2020, also ask how participants resemble those who decline. The framework examines implementation and provides no evidence that a particular AI tool works. [3]

A documentation pilot might cover many encounters performed by very few clinicians. More encounters can provide useful information about those clinicians’ work. They do not create additional independent examples of a colleague learning to use the system. Likewise, recruiting more clinicians from one well-supported department leaves questions about another department’s operating conditions unresolved.

I would ask the team to explain which unit supports each conclusion: a document, an encounter, a clinician or a service. The statistical analysis belongs with people qualified to account for dependence between observations. The leadership task is to avoid letting a large document count imply a broad workforce test.

Make support visible before attributing willingness

Consider two possible extensions of the fictional pilot. One adds clinicians who do comparable work and have the same help available. The other adds a service whose staff cannot attend the introduction sessions or reach support during their working hours. A weaker result in the second extension could have several explanations. Calling its staff resistant would skip the work of distinguishing them.

The 2022 DECIDE-AI guideline asks early clinical decision-support studies to report user recruitment and familiarisation. It is a consensus reporting guideline for a specific evaluation stage; following it does not establish methodological quality or prescribe a study design. [4]

My practical inference is to write down what participation requires and what the organisation supplies. Does learning fit within paid working time? Does access work where the task happens? Can someone obtain help while the task is underway? Record material differences before discussing a group’s willingness. This also helps a supplier understand which apparent product difficulties depend on the delivery arrangement.

Relevant differences should follow the proposed use. For a pharmaceutical information team, they might include access to approved source material and availability of review. For a device manufacturer, they might concern local training and the route for reporting problems. These are constructed examples, with no claim that the same conditions determine outcomes across sectors.

Use the next group to investigate an uncertainty

I would select the next bounded test around an important missing condition. If the proposed expansion includes teams with little immediate support, investigate that condition deliberately, while preserving the support required for safe use. Recruiting another group through the same enthusiastic colleague network may leave the original uncertainty intact.

Purposeful recruitment can expose contrasts worth understanding. It does not automatically produce a representative sample. Nor does random assignment between two tools among volunteers make those volunteers representative of all intended users. An evaluation specialist should help match recruitment and comparison to the conclusion leadership wants to draw.

Comparisons between successive rounds need particular care. Software, training and workload may change while the workforce changes. A different result cannot simply be attributed to the new participants. Keep the relevant conditions visible and describe which explanations the design can distinguish.

Leave unknown experience unknown

A colleague who stops using the tool may have found it unhelpful, encountered an access failure, changed duties or had too few relevant tasks. A missing survey response establishes none of these explanations. I would offer a brief, voluntary way to describe obstacles, separate from individual performance assessment, and report reasons only at a level that protects small groups.

The resulting decision might be to maintain access within the tested scope, fund a missing support condition or commission a further evaluation before expanding. It should identify who remains unobserved. For me, a useful pilot report lets the next team recognise both the opportunity and the conditions attached to it. That gives early success somewhere credible to go.

Sources and further reading

  1. Gupta: first usersY Combinator

    Entrepreneurial stimulus.

  2. Franklin: selected respondentsDiscover Artificial Intelligence

    Volunteers; limited transfer.

  3. RE-AIM: participationRE-AIM

    Settings and staff.

  4. DECIDE-AI: reporting guidelineNature Medicine / UCL Discovery

    Recruitment and familiarisation.

The starting point for this reflection

Gupta: first users

Perspective and interests

This article was developed with AI assistance. The hospital scenario and sector examples are fictional. The proposed approach is an original organisational inference and has not been empirically evaluated.

I am the founder and CEO of aiomics and have a commercial interest in responsible AI adoption in medicine. This article contains no treatment recommendations and does not replace methodological or regulatory assessment.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk