← Ideas & guides

Reflection

The responsibility behind an AI draft

Fluent writing can speed up the preparation of a document. It also changes the task of the person who takes responsibility for its contents.

By Dr. Sven JungmannRetrospective reference date: · Published: · Reviewed:
Watercolour with documents and a magnifying glass, representing the review of source material.

At a glance

  • A usable AI draft needs traceable evidence, visible gaps and clear limits to its claims.
  • Review should focus on consequential statements and omissions.
  • Mapping sources and making targeted changes can support human review; their effectiveness needs checking in the intended setting.

A finished document has an unusual quality: it makes the thinking look complete. The paragraphs are orderly, the terminology fits and the recommendation sounds sensible. Yet a separate task begins for the person who puts their name to it. They need to establish which statements they can actually stand behind.

I am particularly interested in that transition from production to responsibility. As a doctor and the founder of a company developing AI for healthcare, I look at a draft from both perspectives. Fast production is valuable. So is the ability of the responsible person to understand and verify the result with a reasonable amount of effort.

A system can generate a great deal of text. The attention available to the people checking it remains limited. This creates a design problem: the draft should be structured so that consequential errors can be found. A general instruction to provide human oversight does not accomplish that work.

A claim needs a path back to its evidence

Imagine a hypothetical decision brief for a hospital group. It summarizes three publications and recommends trying a new approach. One paragraph says that the studies demonstrate a reduction in workload. On closer reading, one study measured processing time, another measured satisfaction and the third measured how often the system was used.

Each measure can be informative. Together, they do not automatically substantiate the claim in the brief. Someone reading only the polished paragraph must discover this shift in meaning themselves. A well-structured draft would distinguish the different measures and connect each statement to the relevant passage.

In its risk profile for generative AI, the US National Institute of Standards and Technology describes fabricated reasoning and citations. These can make false statements appear more credible. The implication for review is straightforward: a reference must support the claim being made. Its existence alone is insufficient. [1]

For demanding summaries, I would therefore request a small evidence table: the claim, its source, the specific passage and the limits of what can be concluded. This is a working method that needs to prove useful in the setting where it is applied. It seems particularly valuable when several documents are condensed into a recommendation.

Omissions rarely announce themselves

A second difficulty is missing information. A document can contain only accurate sentences and still provide an unsuitable basis for a decision. It might omit a limitation in the study population. It might overlook a conflicting source. It might leave the provisional nature of the figures unclear.

Return to the hypothetical hospital example. If the brief applies findings from a narrowly selected population to the entire hospital group, its meaning changes. A proper review therefore also needs to start from the assignment and the original documents. Someone who only scans the draft sees what made it into the text.

Research on how language models handle long inputs offers another reason to pay attention. In 2024, Nelson Liu and colleagues showed that the position of relevant information could affect performance on the tasks they studied. This does not establish a fixed limit for today's models. It does provide a reason to check separately whether relevant information has been captured. [2]

My recommendation would be to add an explicit check for gaps to the assignment. What information is missing? Which sources disagree? What would need to be known before turning the summary into a recommendation? These questions remain useful even after the model has produced a convincing document.

Review begins before the first sentence

Making approval easier can start with a more precise assignment. What purpose should the document serve? Who will read it? Which source materials are authoritative? What conclusions may the system draw? An internal overview and a document supporting a consequential decision require different levels of care.

Format matters as well. A short draft with clearly defined sections presents a different review task from several pages of continuous prose. A recommendation, its reasoning and the remaining uncertainty should be distinguishable. Readers can then assess which parts they accept and where further clarification is needed.

I would apply another principle to revisions: material that has already been checked should remain as stable as possible. If only one reference needs correcting, regenerating the entire document is often unnecessary. It can introduce new differences and partly invalidate the earlier review. A limited revision with a visible comparison makes renewed approval easier.

This is a practical recommendation about the design of work. Whether it saves time and improves error detection can be tested with actual tasks. The assessment should also include cases where the work needs to stop because the available information cannot support a defensible statement.

Human oversight needs the right working conditions

The World Health Organization's guidance on large multimodal models in healthcare emphasizes well-defined tasks, adequate accuracy and reliability, and the involvement of affected users. For me, this also raises the question of the conditions under which those people will assess an output. [3]

A doctor can review a summary more effectively when the relevant original findings are accessible. An executive can assess a recommendation more effectively when assumptions and uncertainties are visible. An employee is better placed to disagree when a different judgment is explicitly allowed and they do not have to defend their decision to question the AI in the first place.

Additional automated checking can help. A second system might look for contradictions or unsupported claims. Its agreement should not be treated as independent proof when both systems rely on the same inadequate information or similar mistaken assumptions. The quality of the check depends on what it can actually access and compare.

When deciding whether to adopt a system, I would therefore measure two durations: the time to produce a draft and the time until someone can responsibly use it. I would also record the corrections required and the errors missed. If only the first duration improves, the rest of the process deserves a closer look.

For me, the most demanding question to ask of an AI draft is this: can the responsible person see what they are relying on? A document that makes its evidence, limits and unresolved points visible gives them a better basis for judgment. That quality should count from the moment a system is selected and designed.

Sources and further reading

  1. Generative AI risks: fabricated content and citationsNational Institute of Standards and Technology

    Describes fabricated content, reasoning and citations as risks of generative AI.

  2. How language models use information in long inputsTransactions of the Association for Computational Linguistics

    Examines how the position of information in long inputs affected the language models evaluated at the time.

  3. Guidance on the responsible use of large multimodal models in healthcareWorld Health Organization

    Recommends well-defined tasks, appropriate reliability and involvement of affected users when applying AI in healthcare.

Perspective and interests

I am the founder and CEO of aiomics. My professional perspective is shaped by developing and implementing AI in healthcare.

This article was developed with AI assistance. It connects the cited sources with my professional perspective; illustrative situations are identified as examples.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk