← Ideas & guides

Guide

When AI changes: what must be checked before renewed release?

Models, sources and tools change. A practical approach to identifying which prior commitments still hold and who authorises continued AI use.

By Dr. Sven JungmannPublished: · Reviewed:
Two cogwheels lie beside and on a wooden jig with matching recesses.

At a glance

  • Approval concerns a particular working configuration.
  • Stable test cases and legitimately updated source facts need distinct expectations.
  • Technical testing, domain assessment and renewed use need identifiable decision owners.

Imagine a hospital purchasing team using AI to compare approved technical documents. Its application and model name stay the same, but the document search is rebuilt. A qualification previously returned beside a table now arrives separately. The comparison still looks polished. The question for the person responsible is whether the evidence supporting the existing approval still applies.

This example is fictional and concerns internal preparation, with no patient-specific recommendations. I would treat an AI release as an agreement about a particular working configuration. A change should reopen the parts of that agreement it could affect. The practical procedure below is my proposal; it has not been validated as a clinical assessment method.

Identify what has actually changed

The starting point is a public YC Paper Club discussion of how tools, memory and execution environments shape AI agents. It prompted a management question for me: who renews permission to use a system when its surroundings change? [1]

Record the previous and proposed configurations. Include the available model identifier, retrieval sources and their versions, connected tools, relevant resource limits and the human review required before use. A provider’s product name may be insufficient to reconstruct a run. Where precise versions are unavailable, document the limitation and agree how changes will be communicated or detected.

The change might concern software, source material or the work itself. A new supplier-document format can challenge an unchanged application. A changed tool may interpret an empty field differently. Several individually modest changes may interact. The test plan should name those dependencies instead of treating a model upgrade as the only event worth examining.

Separate model performance from the working configuration

Anthropic varied six resource configurations in Terminal-Bench 2.0, holding the Claude model, execution software and tasks constant. Extreme scores differed by six percentage points (p < 0.01). Its February 2026 provider experiment lacks full replication detail and establishes no clinical effect. [2]

For the purchasing team, the implication is a question to investigate: did the new retrieval arrangement preserve the relationship between the table and its qualification? A higher score for the model on an unrelated benchmark cannot answer that. The relevant object is the complete route from an authorised input to the comparison a colleague actually receives.

Link each change to an existing commitment

Before testing, I would write a short change record with four entries: what changes, which existing commitment might be affected, what observation would test it, and who can judge that observation. In the fictional example, the commitment is that a specification remains attached to its conditions. The proposed observation checks whether the accepted comparison retains that relationship after retrieval changes.

A change limited to visual spacing may justify a narrow check if the team can explain why interpretation and subsequent actions remain unaffected. Replacing the model while rebuilding retrieval calls for a broader examination. Where practical, evaluate changes separately before assessing their combination. Otherwise an improvement in one component can conceal deterioration in another.

Keep stable cases and changed facts distinguishable

Use two clearly labelled groups of cases. The first preserves inputs and expected treatment so the previous and proposed configurations can be compared. It includes important behaviours that already worked, known failures and relevant awkward cases. The second covers what the change introduces: a new document structure, additional source types or newly permitted work.

If the underlying source has changed, an old answer may become outdated. Keep the earlier source snapshot and its expected answer together. Have the responsible specialist define the expectation for the new source separately. Otherwise a test can reward obsolete information or report a regression merely because the system correctly follows an updated document.

In the example, one test retains the original table and qualification unchanged. Another introduces a revised document that changes the qualification. The first asks whether meaning survived a technical change; the second asks whether the new information is handled appropriately. Both need records of what was retrieved, what was produced and what the reviewer accepted. Use synthetic or otherwise authorised material.

Test the comparison as well as the application

Anthropic’s January 2026 evaluation guide distinguishes new capabilities from preserved capabilities. It recommends repeated trials, clean test environments and checking whether graders measure the intended outcome. This is engineering guidance from a provider, rather than a trial of organisational or clinical effectiveness. [3]

For our example, keep the evaluation rules and any model used to judge outputs fixed during the comparison. If that judge also changes, rescore some identical outputs with both versions. A changed verdict may originate in the evaluator. A specialist should resolve consequential disagreements against the underlying documents; two agreeing model-generated scores do not settle them.

Report which cases improved, deteriorated or remained unresolved, alongside relevant changes in review work and incomplete attempts. The overall average can conceal a lost qualification in a small but important group. Repeated runs reveal variability; their number needs justification for the decision. There is no universal case count that makes a changed system ready for release.

Name who renews the approval

The technical operator should identify what was tested and which environment will actually be deployed. The responsible specialist should assess whether the affected work remains acceptable. A named person with authority over the use should decide its renewed scope, unresolved conditions and observation period. One person may hold several roles, but the responsibilities should remain distinguishable.

FDA’s August 2025 final guidance for AI-enabled medical-device changes links planned modifications, testing and impact assessment, including interactions between changes. It addresses US device submissions and does not authorise use in Germany. It offers a relevant example of treating change as a lifecycle responsibility; the procedure proposed here establishes no regulatory compliance. [4]

For a clinical application or regulated pharmaceutical process, bring the responsible quality and regulatory functions into that decision. An internal comparison of document summaries cannot supply the evidence required for another intended use. Keep any limitation on the renewed approval visible to the colleagues who rely on it.

Make recovery part of the decision

Before switching, establish whether the previous configuration can actually be restored. An old instruction file will not restore a retired model or a replaced external service. Where recovery is unavailable, define an approved manual alternative or a narrower use that remains acceptable. Record who can interrupt the changed system and how unfinished work will be handled.

After release, observe the specific dependencies the change put at risk. In the fictional example, that means reviewing whether qualifications remain attached to the relevant specifications, including newly arriving document formats. Record findings against the released configuration. The renewed approval should explain what changed, why continued use is justified and what evidence would reopen the decision.

Sources and further reading

  1. YC Paper Club: models and execution environmentsY Combinator

    Starting point for the organisational reflection.

  2. Anthropic: infrastructure effects in coding testsAnthropic

    Six configurations; no clinical evaluation.

  3. Anthropic: evaluating agentsAnthropic

    Provider engineering guidance.

  4. FDA: changes to AI-enabled devicesFDA

    Final US guidance; interacting modifications.

The starting point for this reflection

YC Paper Club: models and execution environments

Perspective and interests

This article was developed with AI assistance. The hospital example is fictional. The proposed testing and approval procedure is my organisational inference; its effectiveness has not been evaluated here. No local pilot data were collected.

I am the founder and CEO of aiomics and have a commercial interest in responsible AI adoption in medicine. This article contains no treatment recommendations and does not replace any required regulatory assessment.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk