At a glance
- A source attribution can change evaluations of identical answers; this establishes no difference in medical quality.
- Newer experiments show that the described task and timing of AI support help shape the response.
- A truthful explanation should make the task, professional review, responsibility and route for questions understandable.
A hospital sends a clear answer. Beneath it appears: “Created with AI assistance.” What does the recipient now know? Software might have corrected the spelling, produced a first draft or suggested medical conclusions. A physician might have checked every sentence or simply initiated sending. The notice leaves several very different processes possible.
For me, the central communication question is what the recipient comes to believe about the work behind the answer. Disclosing AI involvement should help people form an accurate picture. Research shows that a statement about the source can change evaluations. It also gives us reasons to interpret a lower trust rating carefully. Appropriate scepticism can be desirable when someone previously overestimated what an answer could establish.
The same text, three attributed authors
On 25 July 2024, Moritz Reis, Florian Reis and Wilfried Kunde published two preregistered experiments in Nature Medicine. The analysed samples comprised 1,050 and 1,230 people recruited through Prolific. The first was international; the second was matched to the UK population on age, gender and ethnicity. Participants evaluated four short medical question-and-answer scenarios. The attributed source varied between a physician, AI and physician-revised AI. [1]
Within each scenario, the answer remained identical. Every answer had actually been prepared using ChatGPT 3.5 and subsequently edited, supplemented and checked for medical accuracy by a physician. This is the experiment’s strength: differences in the quality of the answer cannot explain the source effect within this design. It examines attributed authorship. It provides no performance comparison between physicians and language models. [1]
Believed AI involvement produced lower ratings of reliability and empathy. In the first study, the standardised reliability difference between human and AI attribution was 0.28, with a 95% confidence interval of 0.13 to 0.43. That describes a small difference on a rating scale. It does not mean 28% less trust or a corresponding change in treatment decisions. [1]
The attribution did not significantly affect comprehensibility in either study. In the second, stated willingness to follow the advice was lower with believed AI involvement. However, when participants could save a link to the supposed platform, neither comparison with the human source showed a significant difference. This limited behavioural measure captured interest in the service. It did not observe actual treatment or sustained use. [1]
What the scepticism leaves unresolved
The design isolates an effect, while leaving its meaning in an individual encounter partly open. Participants read other people’s cases. They could not describe their own symptoms, ask follow-up questions or get to know a treating professional. The second study excluded people who did not correctly recall the attributed source. Its results therefore concern a selected, attentive analysis sample under controlled conditions.
Above all, “AI revised by a physician” describes a collaboration whose practical details remain unclear to the reader. Do people imagine a careful clinical review, a quick scan or merely the possibility of human intervention? A trust rating cannot tell us which interpretation occurred. My inference is that hospitals should initially treat reservations as questions about how clearly their processes are explained. The study provides no justification for concealing AI involvement.
Newer evidence makes the division of work visible
A study by Insa Schaffernak and colleagues, published on 21 August 2026, changes the perspective. In two preregistered experiments, 489 and 570 people evaluated descriptions of ophthalmology consultations. Recruitment used German university participant pools, predominantly employees studying part-time. Data were collected in March–April and June–October 2025. The samples were young, averaging approximately 25 years, and predominantly female. [2]
The experiments compared different AI tasks: additional image highlighting and analyses, or a preliminary diagnostic suggestion. The second study also varied timing. In one condition, the physician consulted AI while making the assessment. In another, the physician first formed and communicated an independent assessment, then reviewed the AI output. Participants evaluated descriptions of these processes; they did not experience actual treatment. [2]
When information was considered concurrently, diagnostic AI support received lower trust ratings than additional visual support. When the physician’s assessment came first, the difference between the two AI types was no longer statistically significant. In the second study, estimated mean trust in the medical decision under diagnostic support was 4.87 with concurrent use and 5.35 with sequential use, on a seven-point scale. That comparison was statistically significant. [2]
This does not establish a clinical recommendation to consult AI only at the end. The study measured evaluations. It did not determine whether the sequence improves diagnosis, changes working time or suits a particular care setting. The process and its verbal description also remain connected. The results do establish a useful boundary for sweeping claims: mentioning AI involvement did not produce a uniform response under these conditions.
Disclosure needs content that can be checked
The ethical case for openness stands independently of which wording wins higher approval ratings. In its principles for AI in health, the World Health Organization connects transparency and intelligibility with human autonomy, accountability and opportunities to question decisions. These are principles to guide design. They do not replace an assessment of the legal requirements for a particular use. [3]
For practical implementation, I suggest explaining four things: what task AI performed; what was subsequently reviewed by a qualified professional; who is responsible for the answer being sent; and how the recipient can raise a question or correction. Each statement must fit the actual process. “Reviewed by a physician” would make a substantial claim if the only checks concerned spelling and formatting.
A fictional example for a service organised accordingly would be: “AI produced the initial draft. Before sending, the responsible physician checked the medical statements and how they apply to your question. You can reach our care team through the reply function.” Every sentence describes work performed or an accessible point of contact. An organisation could use this wording only if it reliably carries out that process.
General health information would require a different explanation. A general professional review cannot support a claim that a particular person’s circumstances have been assessed. The handling of personal data also needs an accurate, understandable explanation. A physician’s approval of the text alone says nothing about where data are stored or who can access them.
Test what people actually understand
A hospital could compare two truthful explanations of the same approved example in a supervised comprehension exercise: a brief notice and a somewhat fuller account of the tasks. This is my own design proposal. The cited studies did not test this particular exercise. The examples should be used outside an ongoing treatment encounter, with the same underlying facts about AI involvement in both versions.
I would then ask people to explain in their own words who checked the medical statements, whether their individual circumstances were considered and whom they could contact about an error. It would also be useful to learn whether they infer a promised service that is actually absent. A high trust rating accompanied by a mistaken belief in individual clinical review would be a warning sign. A more precise and cautious understanding could be a good result.
This distinction also matters for pharmaceutical and medical technology companies, for example when answering professional questions or explaining a product. The experiments discussed here did not investigate those commercial relationships. The extension is my practical inference: an explanation should help recipients judge a specific statement, the scope of its review and the responsible organisation accurately. To do that, the organisation must first know what work was actually performed. It can then describe that work honestly.
Sources and further reading
- Reis and colleagues: Influence of believed AI involvement on the perception of digital medical adviceNature Medicine
Two preregistered experiments with 1,050 and 1,230 participants compare identical answers under different source labels, measuring evaluations, stated willingness to follow advice and limited interest in a platform link.
- Schaffernak and colleagues: Effects of Type and Timing of Clinician-Facing AI Support on Patient Trust in Medical ConsultationsJournal of Medical Internet Research
Two preregistered experiments with 489 and 570 participants recruited through German universities compare ratings of described ophthalmology workflows with different AI tasks and timing; clinical outcomes were not evaluated.
- World Health Organization: Six guiding principles for AI in healthWorld Health Organization
The official principles connect transparency and intelligibility with autonomy, accountability and opportunities to question decisions; they do not establish the effect of a particular disclosure phrase.
Perspective and interests
This article was developed with AI assistance. The organisational examples are fictional. The practical proposals are the author’s inferences from the bounded sources.
I am the founder and CEO of aiomics and have a commercial interest in responsible AI adoption in medicine.



