At a glance
- Time savings become meaningful when review, rework and handovers are included.
- Less pressure, greater capacity and better quality are different goals that require a deliberate choice.
- A useful trial measures improvement alongside any additional burden elsewhere.
Consider a hypothetical example: AI saves a doctor ten minutes when drafting a document. That is welcome in itself. What remains unclear is whether she goes home earlier, has more time for a conversation, or simply starts waiting for the next approval ten minutes sooner. All three outcomes are consistent with the same number.
This distinction interests me as a doctor and the founder of aiomics. AI is often demonstrated at the point where its performance is immediately visible: a text appears, a summary is ready, a question is answered. Its economic and human value emerges over a longer process. It depends on what happens to the result and who still has work to do with it.
That is why I would establish early in any new initiative what the time saved is meant to achieve. It sounds like a small decision. Yet it changes the choice of use case, the definition of success and the expectations of the people who will eventually use the system.
The work continues beyond the draft
Stay with the hypothetical document. Drafting is faster. The doctor still needs to check the information, obtain anything missing and approve the result. Review might become easier because the supporting evidence is clearly mapped to the text. It might take longer because a more elaborate document contains additional claims. The reduction in writing time does not settle that question.
At a minimum, the assessment therefore needs to cover the path to a draft and the path from that draft to a usable result. Handovers count too. Who has to ask follow-up questions? Who makes corrections? How long does work sit waiting? One department can reduce its own processing time while creating additional work for another.
A field experiment by Eleanor Dillon and colleagues offers a useful point of reference. It examined generative AI use among knowledge workers. Changes were most apparent in activities that individuals could adapt independently. Time spent in meetings, by contrast, did not change significantly. This is not evidence about German hospitals. It highlights an important boundary: changing shared working practices involves a different set of conditions. [1]
My practical recommendation for implementation would be to map the complete process on one page. Mark the point where AI is supposed to help and the point where the finished result is actually needed. If those points are far apart, the steps between them deserve particular attention.
Four possible uses for the same time
The first possibility is relief from workload. A working day might involve less unfinished work. A task might require less mental effort. That is a legitimate goal even if the number of completed tasks stays the same. Anyone promising this benefit should subsequently ask about it and check whether employees actually experience it.
The second possibility is additional capacity. For that to materialize, the later stages of the process must be able to absorb it. Preparing more documents has limited value if the same person must approve every decision. The queue could even grow while the first step becomes faster.
The third possibility is better quality. Time released can go into more careful checking, clearer explanations or more thorough preparation. In that case, unchanged completion time could still represent progress. The relevant question is what improvement in quality was intended and how it would become visible.
The fourth possibility is a greater volume of work. As additional analyses and documents become cheap to produce, their number can increase. This may be useful. It can also create more reading and coordination. The organization therefore needs to decide which additional outputs improve a decision and which mainly require more attention.
These possibilities can be combined, but some compete with one another. The same ten minutes cannot be devoted entirely to reducing workload and entirely to additional tasks at once. A leadership team should make the allocation explicit before different groups infer different promises from the initiative.
What I would establish before a trial
First, I would describe the intended result in one sentence. For example: preparing an internal decision should take less time, including review, while maintaining the same level of completeness. That states what needs to be observed. Simply counting the number of drafts produced would tell us too little.
Next, there needs to be a baseline and a manageable period of observation. How long does the process currently take? Which cases are straightforward and which are difficult? What rework is involved? An average can be useful as long as troublesome cases remain visible alongside it. A system that saves a little on most tasks but creates substantial extra work on a few deserves closer examination.
The Institute for Healthcare Improvement distinguishes outcome, process and balancing measures. Balancing measures look for unwanted consequences elsewhere. Applied to our example, I would examine total effort, actual use and the burden of corrections together. This application is my proposal for assessing an AI implementation. [2]
I would also specify when the trial should lead to a decision. Who will assess the result? What improvement would justify wider use? Which recurring problems must be resolved first? Without this agreement, a trial can become a permanent interim arrangement whose value nobody evaluates clearly.
Our own impressions need checking too
METR's research illustrates the care needed with broad productivity claims. In an experiment involving experienced developers and tools available in early 2025, work took longer with AI. The study had a narrow scope. In a February 2026 update, METR described changed conditions and substantial selection problems that made its newer measurements difficult to interpret. [3] [4]
I take this as a reason to keep checking the effect in our own setting. Tools improve, people learn and tasks change. An earlier finding can become outdated. Equally, an enthusiastic first impression can overlook work that appears later.
The essential leadership task therefore begins with agreement about the benefit. If the goal is relief, people should feel it in their working day. If the goal is greater capacity, it needs to reach the end of the process. If the goal is higher quality, there needs to be a clear understanding of what will become more reliable.
Before the next implementation, I would ask the people involved to complete this sentence together: when AI gives us time back, we will use it for this. Their different answers may already reveal the decision that is still missing.
Sources and further reading
- How generative AI changes work patternsMicrosoft Research
A field experiment on changes in individual work patterns with generative AI, with limited transferability to other working environments.
- Outcome, process and balancing measures for improvementInstitute for Healthcare Improvement
Explains outcome, process and balancing measures for assessing improvements and possible consequences elsewhere.
- An experiment on experienced developers' productivity with early-2025 AIMETR
A randomized study of experienced developers' completion time using early-2025 AI tools in familiar projects.
- Why METR is changing its productivity experimentMETR
Describes selection and measurement problems in the later experiment that limit interpretation of its newer productivity estimates.
Perspective and interests
I am the founder and CEO of aiomics. My professional perspective is shaped by developing and implementing AI in healthcare.
This article was developed with AI assistance. It connects the cited sources with my professional perspective; illustrative situations are identified as examples.


