← Ideas & guides

Guide

Maintaining shared AI instructions: test changes and preserve versions

A practical maintenance agreement for shared AI instructions: purpose, test cases, changes, approval and returning to an earlier version.

By Dr. Sven JungmannRetrospective reference date: · Published: · Reviewed:
Interlocking index cards form a supporting bridge.

At a glance

  • Assess a wording change against the desired behaviour.
  • A tested version includes the instruction, environment, examples and an accessible owner.
  • An earlier text version cannot fully restore the behaviour of a model that has since been replaced.

A shared AI instruction can change quickly. Someone adds a phrase, a colleague removes a paragraph, and a new model version becomes available. Each change may look reasonable. A few weeks later, it can still be difficult to explain why the same assignment produces different results. For me, maintaining an instruction therefore begins with a simple question: which version produced which behaviour, under which conditions?

This guide concerns instructions for bounded organisational tasks. Its example is fictional: a medical technology company uses a template to draft an internal action list from approved meeting notes. Personal or confidential information may be processed only within the environment actually authorised for it. A well-written template does not extend that permission. The procedure described here also provides no approval for patient-specific decisions.

Why an improvement in wording deserves a check

In their 2024 paper, Sclar and colleagues examined, among other experiments, 53 classification and multiple-choice tasks using language models of that period. Formatting changes intended to preserve meaning could substantially change measured accuracy; favourable formats transferred only weakly between models. These experiments establish no error rate for current clinical applications. They support a narrower question: does the desired behaviour survive a change? [1]

The voluntary NIST profile for generative AI, published in July 2024, includes traceable versions and testing user instructions among possible measures. It does not evaluate a particular workplace template collection. The maintenance agreement below is my own proposal for applying these ideas to a shared instruction with a manageable amount of effort. [2]

Define the assignment first

The example template should collect agreed actions, explicitly assigned responsibilities and existing deadlines from meeting notes. It should identify open questions. A suggestion made during discussion should retain its provisional status. If nobody has accepted an action, responsibility remains unassigned. This creates an assignment that can be checked and used to assess a change.

Six short entries sit alongside the instruction: permitted purpose, suitable inputs, excluded uses, responsible person, tested environment and current version. The environment includes the model, application and available additional functions wherever that information is accessible. Record the test date too. If a provider conceals the precise model version, document that uncertainty. An invented version number would be worthless during a later investigation.

Maintenance needs an accessible person with responsibility for the substance. In this example, the person responsible for project coordination can assess whether decisions and discussion points remain distinguishable. A technical colleague can manage distribution. In a small company, one person may hold both responsibilities. People should still be able to identify who approves a changed version for shared use.

Build a small collection of test cases

I would start with a few fictional meeting records whose expected treatment is clearly described. One contains straightforward actions. A second contains a suggested deadline that nobody has agreed. A third assigns conflicting responsibilities to two people. A fourth records a later correction to an earlier decision. A fifth provides an empty or incomplete input. This number is an organisational starting point, with no claim to be a scientifically established minimum.

Define the expected treatment before running the test. For the proposed deadline, the criterion might be: identify it as a suggestion and do not present it as an agreed due date. For the correction, the result should make clear which statement has been superseded. Judging stylistic elegance would obscure these differences. Preserved meaning, visible uncertainty and usable attribution matter more for this task.

Keep some cases aside for a later check. Otherwise a template can be repeatedly improved against the same examples until it mainly reflects their peculiarities. Ask different colleagues to try the input instructions as well. If real information is considered for testing, permission to use and retain it must be established separately. Fictional material is sufficient for this first maintenance exercise.

Treat a change as a testable proposal

Suppose the action list is too long. A colleague proposes reducing each action to one sentence. The change note could say: the new version should reduce repetition while preserving qualifications and missing ownership. This states the improvement being sought and the property that shortening might put at risk.

Give the old and new versions the same permitted examples under conditions that are as similar as possible. Repeated runs can help establish whether an observed difference recurs. Their number depends on the importance and variability of the task; a few successful answers provide no general reliability guarantee. A short note should record observed differences and the decision about whether further examination is needed.

In the fictional example, shortening removes the qualification that a deadline still requires agreement. That version should therefore be revised. Retain the change note so the next person can understand the attempt. A rejected change contributes to the organisation's knowledge about the template. It may explain why an apparently cumbersome phrase has been necessary.

Connect approval, recovery and continued use

Give the released version a distinct identifier and a reference to its predecessor. Anyone copying a template should be able to find the current shared version. A brief notice identifying superseded versions helps people interpret older copies. If several variants are needed in parallel, each should state its own purpose. An undisclosed adaptation for one department otherwise makes subsequent comparisons harder.

Prepare for returning to the previous version too. This includes its instruction and known environment, insofar as that environment remains available. If the provider has replaced the underlying model, an old text template cannot fully restore the previous behaviour. A fresh assessment or agreed manual procedure is then needed. The saved version remains useful as a basis for comparison.

A new model version, changed inputs or a recurring misunderstanding can trigger an additional check. A scheduled review complements these triggers. The organisation chooses its interval according to frequency of use and the consequences of an error. An occasional overview and an action list processed further every day deserve different levels of attention.

What a maintained collection makes visible

A useful collection answers the same practical questions for each template. What is it intended for? Which inputs have been tested? What difficulties are known? Who receives feedback? Which version is in shared use? These details help colleagues select an appropriate template and recognise an unsuitable use before proceeding.

For the next internal meeting, I would select a frequently copied instruction and reconstruct its most recent change. If nobody can explain why a paragraph was added, that provides a useful starting point for maintenance. The result should be a comprehensible, testable version with an accessible owner. Its value lies in allowing the organisation to develop its way of working deliberately.

Sources and further reading

  1. Sclar and colleagues: sensitivity to instruction formattingICLR / University of Washington

    The conference version examines, among other experiments, 53 selection and classification tasks using models of that period. It tests no clinical application.

  2. NIST: versions and testing of user instructionsNIST

    The voluntary July 2024 profile addresses traceable versions and testing user instructions; the workplace maintenance agreement is an original inference.

Perspective and interests

This article was developed with AI assistance. The organisational examples are fictional. The practical proposals are original inferences from the sources within their stated limits.

I am the founder and CEO of aiomics and have a commercial interest in the responsible adoption of AI in medicine.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk