← Ideas & guides

Guide

When may AI-assisted literature screening stop?

Finding few new relevant records is insufficient justification. How specialist teams and leaders can preserve what was assessed, what remains open and which decision a briefing may support.

By Dr. Sven JungmannPublished: · Reviewed:
A row of upright paper folios on a library table continues behind a translucent blue-grey screen.

At a glance

  • Preserve the states retained, excluded after assessment and unscreened through the handover.
  • Label a briefing closed because of a deadline as provisional and limit its use.
  • Document why screening ended, along with the next assessment step and its owner.

A literature briefing for the leadership team is due on Friday. AI has moved promising publications to the top of the queue. Few relevant records have appeared for a while. May the team stop? To make that decision, I would first establish what claim the completed briefing is supposed to support and what will happen to the unread records.

For medical departments, pharmaceutical companies and medical technology developers, the same literature can serve different purposes: initial orientation, preparation for an investment or a systematic assessment. A tight deadline therefore belongs in the assignment. It should remain visible in the description of the result when the work changes hands.

What the research adds

In 2026, Kempny and colleagues examined five previously labelled datasets using ASReview 1.6.1. Three simple stopping criteria did not consistently recover every relevant record in the simulations. The study tested complete retrieval under those conditions, without comparing current language models or measuring clinical outcomes. [1]

In a 2024 methodological commentary, Callaghan and colleagues call for stopping rules with a target recall, reliable uncertainty estimates and evaluation across datasets. This is expert guidance; it supplies no universal number of unread publications that an organisation can safely leave behind. [2]

My organisational proposal is to give every briefing a visible record of its completion status. The fields and examples below are my own working aid. They do not replace a study protocol or expert specification of a statistical stopping rule.

Preserve three states through the handover

Imagine a fictional search containing 10,000 records after cleaning. The team has examined 1,000 titles and abstracts. Of those, 80 have been retained for further assessment and 920 excluded against the criteria. The last 60 records examined were ineligible. Another 9,000 remain unscreened. These numbers are invented.

The briefing should preserve those three states separately: retained, excluded after assessment and unscreened. Nor do the 80 retained records yet represent 80 finally included studies. Several publications may concern the same study; in this example, full-text assessment is still pending.

Even completing the retrieved collection does not establish that the original search found every relevant publication.

Dividing 80 by 1,000 gives an eight per cent retention rate among the records examined so far. It does not establish the proportion of all relevant records that has been found. The missing quantity is the number of relevant records in the entire collection. Dividing 80 by 80 would also be misleading here: the denominator would contain only records already found.

I would therefore keep the unscreened collection explicitly visible at handover. A hidden remainder might later be mistaken for completed assessment. A visible remainder can receive a follow-up assignment, with an owner, a date and a clear intended use.

Put the decision question before the reading list

For a preliminary executive briefing, the question might be: which assumptions behind our proposal should we investigate next? An initial selection can already help, provided its limits remain visible. I would withhold a much broader statement such as “the literature shows no relevant disadvantage” on the basis of that same interim result. The example establishes only what has been examined so far.

I would specify three things in the assignment: the decision, the kind of claim permitted and the treatment of gaps. For example: “Preliminary orientation for planning further assessment; no final evaluation; unscreened publications will be listed separately.” Someone who did not participate in the research can then understand its status.

For a systematic assessment, the responsible specialist team must establish the applicable methods. My proposal concerns the handover between research and decision-making. It does not turn limited screening into a complete assessment.

Give stopping its own justification

I would test the statement “we are stopping because the latest records were irrelevant” against a short decision note. It should state the volume examined, the remaining collection, the method applied and the responsible person. If the deadline alone determines the decision, label the result “provisionally closed because of the deadline”. If a justified stopping rule has been applied, document its name, assumptions and assessment.

The AI-generated ordering also belongs in that note. A sequence of low-ranked records is not automatically a random sample of the remainder. Anyone wishing to use a sample as justification needs an appropriate sampling and analysis plan. Reading a handful of additional texts should not retrospectively acquire the label of statistical assurance.

In the discussion, I would ask one further question: what unresolved finding could change the decision ahead? That can create a focused follow-up assignment. Such a search supplements the existing work; finding nothing does not, by itself, establish completeness.

A handover that remains understandable

A useful accompanying note might read: “As of: [date]. Purpose: [decision]. Examined: [count]. Retained for further assessment: [count]. Excluded after assessment: [count]. Unscreened: [count]. Stopped because of: [deadline or justified rule]. Permitted use: [scope]. Next assessment step and owner: [details].”

Also preserve the search date, search scope, eligibility criteria, software version and associated decisions so the specialist team can resume the work. When updating it, make clear which records have been added and whether the criteria have changed. Earlier judgments then need to be checked for continuing applicability.

As the founder of aiomics, I am particularly interested in which responsibilities remain visible when a tool’s output becomes an organisational decision. A leader should be able to tell whether the document provides initial orientation or a completed assessment. For me, that is part of the quality of the result, alongside the selection of texts itself.

Sources and further reading

  1. Kempny (2026)BMC Medical Research Methodology

    Simulation and limits.

  2. Callaghan et al.: methodological commentary (2024)Systematic Reviews

    Methodological requirements.

The starting point for this reflection

Kempny (2026)

Perspective and interests

This article was developed with AI assistance. The numerical example is fictional. The handover note is an original organisational proposal, not a validated statistical stopping rule or a complete systematic review procedure.

I am the founder and CEO of aiomics. This article reports no customer-project results. Sources were checked on 29 September 2026.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk