← Ideas & guides

Reflection

Ambient documentation: what remains after working hours?

What clinical studies show about note-writing time, exhaustion and after-hours work, and which measurements hospital medical leaders need for a local pilot.

By Dr. Sven JungmannRetrospective reference date: · Published: · Reviewed:
A desk lamp illuminates the end of a long paper ribbon folded into a compact stack.

At a glance

  • Documentation time, experienced burden and after-hours work are separate outcomes.
  • Activity logs can miss editing in other applications.
  • A pilot needs an explicit time boundary, a baseline comparison and its own quality assessment.

A hospital introduces software that listens to consultations and drafts the documentation. In a fictional evaluation meeting, clinicians report that the day feels less fragmented. The electronic record shows less time spent writing notes. Yet some people still finish their documentation at home. All three observations can be accurate. For a medical director, the next decision depends on which improvement the organisation promised and whether its measurement captures the work that remains.

Ambient documentation creates several plausible benefits. It may reduce typing, make a consultation easier to follow or help people finish within their working hours. Each requires its own evidence. Recent clinical studies provide reasons to test these systems and reasons to measure after-hours work explicitly. They do not provide a transferable promise of an earlier finish for a German hospital department.

What the randomized trials actually measured

Lukac and colleagues randomized 238 outpatient physicians at UCLA to two ambient documentation products or a control group. The trial ran from November 2024 to January 2025 and was published in November 2025. In the second month, Nabla reduced recorded time per note by 9.5% relative to control; the DAX estimate was 1.7% lower, with uncertainty compatible with either a reduction or an increase. Neither intervention demonstrated a reduction in electronic-record work outside scheduled hours. [1]

The study also reported favourable signals on measures of work experience. These were secondary outcomes requiring cautious interpretation. A particularly relevant limitation was discovered during the trial: the record system did not capture time spent editing within the vendors’ platforms. The trial was also registered after it had begun. English-only consultations, short follow-up and use in about one-third of recorded visits further constrain interpretation. Those details matter when a hospital translates the result into a procurement expectation. [1]

Afshar and colleagues offer a different view. Their 24-week randomized trial introduced Abridge in stages to 66 willing practitioners in US outpatient clinics. Work exhaustion and interpersonal disengagement improved. Professional fulfillment did not meet the prespecified statistical threshold. Recorded note-writing time fell by about 22 minutes per eight hours of patient-care time. The estimate for electronic-record work after hours was sensitive to extreme observations: after excluding the highest 3%, its confidence interval included no change. [2]

This leaves a practical distinction. There is evidence of benefit on an exhaustion measure alongside less secure evidence about after-hours work. A favourable questionnaire result can be valuable in its own right. It cannot supply missing clock-time evidence. The open-label design and recruitment of willing users also matter when judging how far the result might extend to colleagues who are less eager to adopt the software. [2]

Larger field studies preserve the distinction

Rotenstein and colleagues studied 8,581 outpatient clinicians at five US academic institutions, including 1,809 who received access to an AI scribe. Their April 2026 observational comparison associated receipt of access with 16 fewer documentation minutes per eight scheduled patient-care hours. The main analysis found no statistically significant reduction in electronic-record work outside scheduled hours; some sensitivity analyses differed. Access was largely based on voluntary uptake, so unmeasured differences remain a possible explanation for part of the association. [3]

A May 2026 study by Husa and colleagues analysed 1,547 clinicians who had used ambient documentation for at least 25 encounters in one month. Its interrupted time series found no immediate reduction in documentation after hours, alongside a gradual downward change in the subsequent trend. Selection of active users and the retrospective design limit causal conclusions. It adds a reason to examine trajectories over time, without settling what another department should expect. [4]

Taken together, these results support separate decisions about documentation time, experienced burden and work beyond scheduled hours. A statistically uncertain estimate leaves a range of plausible effects. Equally, an average improvement may conceal people whose evenings remain difficult. The studies mostly concern US outpatient practice. Inpatient handovers, multilingual conversations, different documentation obligations and different record systems require local investigation.

Start with the boundary of the working day

I would include after-hours work as an explicit outcome whenever relief outside working hours is part of the pilot’s rationale. The following is my proposed evaluation approach; the cited studies have not tested it. Its first task is to agree what counts as after-hours work. Time outside booked consultations can include paid documentation sessions. Time after a person’s agreed working hours answers a different question.

Before deployment, the clinical lead, participating staff and evaluation lead should agree which schedule defines the boundary, how leave and on-call work are handled, and how changed appointment templates will be recorded. Keep both the original definition and a record of later changes. Extending the recorded schedule could reduce a measured overrun while leaving the person working just as long. That would need an explanation in the result.

Use a short baseline diary alongside authorised activity logs. On a small set of agreed observation days, participants record when documentation finishes, which application they use and whether unfinished work resumes later. Include representative busy and quieter days. The diary should also permit a brief reason for a late finish, such as an interruption or a backlog unrelated to note drafting. It can reveal omissions in the logs, while its own reporting burden and imperfect recall must remain visible.

Follow the work across applications

The UCLA measurement limitation suggests a concrete check before trusting a time reduction. Map where a draft is read, edited, transferred and approved. Establish which of those steps each data source captures. A fall in the record system’s note-writing time becomes more informative when the evaluation also knows whether editing has moved into another application.

In a constructed example, a physician spends less time in the record editor but now reviews the draft on a phone between consultations. If the evaluation only measures the editor, it cannot estimate the whole change. A second possibility is that documentation finishes earlier but unrelated inbox work extends into the evening. The pilot should retain the distinction between documentation after hours and all work after hours. Their divergence can identify a further organisational question.

Agree the permissible data collection, information and consent arrangements locally before recording conversations or observing staff activity. Use approved systems and aggregated reporting with safeguards for small groups. A management report needs patterns of work and correction burden; identifiable conversation content should stay within its authorised purpose. Keep the evaluation separate from individual productivity ranking, and make its purpose and access arrangements clear to participants.

Pair time with experience and quality

Repeat an appropriate validated burden measure, using the same instrument and timing at baseline and follow-up. Keep exhaustion, professional fulfillment and perceived time relief as distinct results. Add a short explanation of what changed in the consultation or the evening. These accounts can help interpret an apparent mismatch between logs and experience. They do not establish the cause of that mismatch by themselves.

Quality needs its own assessment. Agree how appropriately qualified reviewers will check a permitted sample for omissions, incorrect attribution and required corrections, and record the work involved. Afshar’s note-quality assessment used a language model to judge unedited AI notes against transcripts. That particular assessment does not establish patient-outcome equivalence. The pilot’s time findings should therefore be reported alongside the scope and limitations of its quality checks. [2]

Set the observation period, comparison and meaningful improvement before viewing the results. Separate initial learning from later use, record software changes, and include non-users and discontinued use in the main implementation picture. A concurrent comparison or randomized rollout, where feasible, can strengthen inference. Show the distribution of outcomes and explain missing observations. An average alone cannot show whether a heavily burdened subgroup benefits.

What the decision should say

The final record can state separately whether documentation became faster, whether after-hours work fell and whether reported burden improved. It should identify what remained unmeasured and whether the quality checks support continuation. If the agreed after-hours objective remains uncertain, specify the additional observation needed and a decision date. The organisation can recognise a useful improvement while remaining precise about the promise it has yet to demonstrate.

Sources and further reading

  1. Lukac and colleagues: randomized trial of ambient AI scribesNEJM AI

    238 outpatient physicians; different results for the two products; vendor-platform editing omitted from the time measure and registration after trial commencement.

  2. Afshar and colleagues: randomized rollout and professional well-beingNEJM AI

    66 practitioners over 24 weeks; exhaustion and note time; professional fulfillment below the required statistical threshold; after-hours sensitivity analysis; model-based note-quality assessment.

  3. Rotenstein and colleagues: time expenditure across five academic institutionsJAMA

    Observational comparison of 8,581 clinicians, including 1,809 receiving access; time normalized to eight scheduled patient-care hours and uncertainty about after-hours work.

  4. Husa and colleagues: documentation burden and trends after introductionJAMA Network Open

    Retrospective interrupted time series of 1,547 active users; no immediate reduction in documentation after hours but a change in the subsequent time trend.

Perspective and interests

This article was developed with AI assistance. The meeting and organisational examples are fictional. The proposed evaluation approach is the author’s inference and was not tested in the cited studies. No local pilot data were collected.

I am the founder and CEO of aiomics and have a commercial interest in responsible AI adoption in medicine. This article contains no treatment recommendations.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk