At a glance
- Every percentage needs a clearly described unit of work.
- The comparator, error consequences and subsequent work determine value.
- The next investment should fit the claim actually investigated.
The official TrialGPT project page presents a specific number: 42.6% less time spent screening potential trial matches. [1] That is an interesting prospect for a pharmaceutical company. As soon as the number is used for staffing, a contractual target or a recruitment forecast, its scope needs to be described precisely.
I use this publicly verifiable claim to work through five questions for pharmaceutical leaders. The purpose is to assess a published statement and the investment decision it could inform. The proposals below are my analytical interpretation. They do not constitute an assessment of a product currently on sale or advice on individual patients' clinical eligibility.
On 18 November 2024, the National Library of Medicine described a small pilot: two physicians reviewed six patient summaries against six clinical trials, with and without assistance. Its announcement rounded the time reduction to 40%. [2] The project page, announcement and research article concern the same undertaking. They are not three independent confirmations.
The unreviewed TrialGPT 2.0 preprint of 1 September 2026 adds 288 retrospective cases and 27 prospective tumour-board cases. The latter yielded additional clinician-selected trial options; actual enrolment was not assessed. [4] My worked example remains the 2024 claim. A procurement decision today should also examine the newer system version.
1. What work was actually measured?
In the research article, the decision was made after reading the case: definitely ineligible or worth further investigation. The six partly synthetic vignettes concerned oncology trials. [3] This gives the investigated step a comprehensible description. A leader can then establish where that step occurs in their own process and who performs it.
I would first turn the claim into a complete sentence: under the described conditions, a particular preliminary decision was made faster. The organisation should then add the other steps between receiving an enquiry and actual participation in a trial. These might include obtaining missing information, contacting the trial site, further examinations and the person's own decision about participation.
The economic significance depends on how much of the organisation's work the measured step represents. An internal calculation would therefore need to establish the current time spent on that step separately. It could then investigate how much of a potential saving reaches the complete process. The published percentage alone does not provide that local baseline.
2. How similar is our starting point?
A medical leader could address this question by describing ten typical enquiries from the existing process, initially without personal information. What documents arrive? How complete are they? Who already knows the clinical history? Which information needs to be obtained from other systems? This description reveals the work that would precede AI assessment in the proposed use.
Comparability has several dimensions. Two institutions may work in the same indication while receiving very different inputs. A structured case overview creates different demands from a collection of lengthy documents. Reviewing a selected set of trials is a different search problem from identifying all potentially relevant opportunities. A proposed transfer should name these differences individually.
I would pay particular attention to the origin of missing information. Could a medical professional resolve the gap from the authorised materials? Is another enquiry necessary? Can the decision remain open until an answer arrives? This produces an examinable requirement for the process. It helps an organisation avoid silently treating a favourable test setup as its normal clinical situation.
3. What is the comparator, and which errors matter?
For its own decision, leadership needs a concrete comparison. This might mean today's screening by the same team with the same information available. It could also mean an already improved process without the proposed AI. The important question is the additional contribution expected from the change relative to a realistic alternative.
The errors deserving particular attention should also be established in advance. An opportunity wrongly excluded and an unsuitable opportunity unnecessarily pursued have different consequences. A single average accuracy measure can combine both and conceal a shift that matters to the organisation. I would ask for both error types to be reported separately, alongside unresolved cases and the work needed to clarify them.
The research article reported no statistically significant accuracy difference in the small pilot. [3] That does not establish a general assurance of equivalent quality. An organisation wishing to set a minimum quality requirement for its own use must define it explicitly and assess it with a suitable study size. That is a question for the responsible clinical and methodological team.
4. What work and decisions follow?
Consider a fictional situation: a pharmaceutical company receives a list of possible contacts sooner through an improved initial assessment. Trial sites can still answer follow-up enquiries only at fixed times. The list arrives earlier; the response time of the following step stays the same. What value does the improvement provide under these conditions?
There could still be value. Employees might plan their work earlier or examine more enquiries carefully. Additional demand might also emerge that other participants must handle. The intended consequences belong in the decision paper. An unchanged recruitment duration would not automatically make bounded relief from work worthless. Equally, reduced workload cannot simply be converted into additional trial participants.
For a German or European application, I would also explicitly provide for checking practical accessibility, current availability of places and the responsible contacts. This is a proposal for designing the process. A computationally plausible opportunity becomes useful to the person concerned when it leads to a dependable next action, with its remaining uncertainty clearly communicated.
5. Which next investment does the claim justify?
For me, a publication of this kind first justifies investigating a suitable local question more closely. The decision could be to fund a bounded feasibility assessment for a defined enquiry route. The organisation would name the inputs, the responsible function and the outcome it wants to judge afterwards. Its scope and use of information must fit the local authorisations.
The decision paper should also state which assumption would alter a larger investment. If most existing work occurs before the initial eligibility review, improving data availability could take priority. If that review constitutes a meaningful constraint, a more extensive investigation could be justified. Either answer could emerge from the same initial interest in the research.
A contractual target should correspond to the outcome that can be influenced. An assessment of a screening step can specify its processing time, quality and need for rework. A target for complete recruitment also requires the conditions affecting other participants. This distinction makes responsibility for commitments more comprehensible and gives procurement, medical leadership and executives a shared language.
Make the claim usable in a meeting
I would expect five brief answers on one page: the work studied, similarity to our use, comparator and error consequences, subsequent work, and the next justified investment. An unanswered point remains an open question. The research achievement can then stay visible while its economic significance is examined concretely, with the additional assumptions available for discussion.
As a physician and founder, I am interested in this translation between research and decisions. A good scientific number can improve a choice when people explain what it means and which assumption they are adding. That ability is particularly useful in pharmaceutical leadership: it connects interest in research with precise commitments to teams, trial sites and, ultimately, the people for whom a study could be relevant.
Sources and further reading
- TrialGPT project pageNLM / NIH
Public screening-time claim.
- NLM announcement about TrialGPTNLM / NIH
Description of the small pilot.
- TrialGPT (2024)Nature Communications
Task and accuracy comparison.
- TrialGPT 2.0 preprintarXiv
New cohorts and bounded outcomes.
Perspective and interests
This article was developed with AI assistance. The organisational examples are fictional. The practical proposals are the author’s inferences from the bounded sources.
I am the founder and CEO of aiomics and have a commercial interest in responsible AI adoption in medicine.



