← Ideas & guides

Reflection

What does a second AI model actually check?

Two models agree. Which errors still survive? A study and a fictional calculation help leaders assess an additional AI review step before buying it.

By Dr. Sven JungmannPublished: · Reviewed:
Two mirrors reflect a pear while a small blue stone behind it remains outside their reflections.

At a glance

  • The value of a second AI reviewer can be assessed through the errors it catches and passes.
  • The share of errors missed and the error rate among accepted documents have different denominators.
  • A useful comparison also records unnecessary flags, unresolved cases and review effort.

A team adds a second AI model to check the first model’s work. One writes a briefing, the other approves it. The purchase decision now depends on a deceptively simple question: what evidence shows that the extra step catches consequential errors?

Imagine an internal investment briefing that confuses a supplier’s current capability with a feature promised for next year. Two fluent accounts of the same mistaken assumption leave the decision exposed. This is a hypothetical example. The useful question for a leadership team is how to test the checking arrangement before relying on its approval.

What the research actually measures

Kim and colleagues’ ICML 2025 study analysed 71 models in its HELM dataset. Among questions both models answered incorrectly, pairs selected the same wrong option about 60% of the time on average. This conditional result from multiple-choice tasks does not estimate errors in an operational document-review workflow or today’s models.[1]

For a purchasing decision, I would treat agreement as an observation requiring further investigation. A second answer can add useful information. Its value as a control depends on which errors it catches and which it allows through. Counting approvals alone cannot answer that question.

Start with the documents that pass

Consider a deliberately fictional test of 100 briefings. Qualified assessors, using the underlying records, establish that 20 contain a consequential factual error. The second model flags 15 of those erroneous briefings and also flags five correct ones. It passes the other 80.

Five incorrect briefings have therefore passed. They represent 25% of the original errors, because five divided by 20 is one quarter. Among the 80 accepted briefings, the error rate is 6.25%, because five divided by 80 is 0.0625. These percentages answer different questions. Neither figure comes from the study.

The checking step has caught useful errors in this example. It has also sent 20 briefings for additional attention, including five that were correct. Whether this is worthwhile depends on the consequences of the remaining errors and the work required to resolve the flags. We have measured neither time saved nor the quality of any subsequent correction.

This arithmetic also explains why a large approval percentage can be misleading. A system that approves everything passes 100% of cases while catching no errors. The relevant evidence includes what survives its approval.

Write a specification for the checking step

NIST’s voluntary AI Risk Management Framework 1.0 calls for assessing controls and testing under conditions resembling deployment. Its measurement guidance also asks organisations to document evaluation methods and their limits. That is a useful basis for evaluating the additional review step.[2]

The following is my proposed application of that principle. It is an evaluation design, with no claim that this particular procedure has been validated.

First, define the error that would change a decision. For an investment briefing, that might be an unsupported revenue figure, a misdated capability or an omitted qualification. Separate consequential errors from stylistic disagreements. Assign a named person responsibility for the definition.

Second, assemble cases whose reference answers can be established from authoritative records by suitable assessors. Record unresolved disagreements instead of forcing them into a correct-or-incorrect label. Include ordinary cases and plausible exceptions, such as conflicting document versions. Document how closely this set resembles the intended work.

Third, test the actual arrangement. Keep the model versions, instructions and available sources fixed. Specify whether the checker sees the first answer immediately or examines the evidence first. Compare results on the same cases with and without the additional control. Changing that arrangement creates a different test.

Fourth, report the counts behind the percentages: erroneous documents caught, erroneous documents passed, correct documents flagged and cases left unresolved. Show the accepted share, error severity and review effort alongside them. A small test with no observed failures leaves uncertainty about rare failures; it provides no universal safety certificate.

Finally, retain fresh cases for evaluation after development. Recheck when models, documents or tasks change. A successful test supports a bounded decision about the tested arrangement.

As the founder of aiomics, I find this a useful way to frame the management question: who defines an acceptable residual error, and what evidence will they examine? The next procurement discussion can ask for a case-level comparison showing what the second reviewer caught, what it missed and what further work its flags created.

Sources and further reading

  1. Kim: model errorsPMLR

    Conditional agreement.

  2. NIST: risk management, version 1.0NIST

    Evaluation of controls.

The starting point for this reflection

Kim: model errors

Perspective and interests

This article was developed with AI assistance. Both company examples and all figures in the calculation are fictional. The proposed evaluation procedure is an original methodological inference; its effectiveness has not been evaluated here.

I am the founder and CEO of aiomics and have a commercial interest in responsible AI adoption in medicine.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk