At a glance
- Specify whether the simulation should prepare questions or estimate customer preferences.
- Require a comparison with relevant human responses reserved for evaluation.
- Keep generated answers, observed people and the authorised decision distinct in every presentation of results.
A fictional medical technology company is choosing the next feature for its customer portal. One option shows delivery updates. Another lets customers download service documents. The team asks an AI model to answer as purchasing managers, generates a thousand responses and finds a clear favourite. It now wants to commit the development budget.
Before that commitment, I would ask what the exercise has established about the people who will use the portal. A thousand generated answers describe the output of the chosen procedure. The commercial decision concerns customers with actual responsibilities, existing alternatives and limited time. Connecting those two requires evidence.
As the founder of aiomics, I find the underlying allocation question useful: what can we learn cheaply before committing development time, and which uncertainty still requires contact with the intended users? Simulated conversations can be a proposed way to explore questions. Their status must remain visible when an exploratory exercise reaches a budget meeting.
Decide what the simulation is allowed to support
For the fictional portal, I would separate three possible uses. The team could generate possible objections to each feature, rehearse an interview or estimate which feature customers prefer. These uses demand different evidence. An imagined objection can become a question to investigate. An estimate presented as customer preference needs a defensible connection to those customers.
Suppose the simulated purchasing manager says that delivery updates matter because colleagues keep calling about delayed orders. That suggests an interview question: when was the last such interruption, what information was missing, and how was it obtained? It does not establish how often this happens across the customer base. The useful next step is to investigate the proposed mechanism.
I would label the exercise accordingly: generated hypotheses for customer interviews. If the team later wants a percentage in a board paper, it must justify that additional use. Repeating generation more often does not by itself supply new observations of customers.
Read the research at the level of its task
Argyle and colleagues used GPT-3, conditioned on real respondents’ backgrounds, to reproduce selected patterns in US political surveys. Their 2023 study makes fidelity dependent on the context and population. It establishes no accuracy for our portal decision. [1]
Bisbee and colleagues compared ChatGPT 3.5 Turbo responses with 7,530 profiles from 2016 and 2020 US political surveys. Their 2024 paper reports similar overall averages alongside narrower variation and distorted relationships. Its historical political setting does not measure current customer-research products. [2]
These studies differ in models, questions and procedures. They leave the buyer with a practical question: what evaluation supports the precise use being sold? A demonstration of fluent dialogue provides a chance to inspect the dialogue. The intended commercial claim still needs its own test.
Ask for a comparison that could change the decision
For the portal example, I propose a short evidence brief with five entries. This is my own organisational method, without a demonstrated effect on decision quality. A research specialist should help design the actual validation and determine the sample size.
First, specify the decision. Are we choosing which feature to investigate, which prototype to test or which feature to build? Record the cost of being wrong and the point at which the choice becomes difficult to reverse. An exploratory discussion and a substantial development commitment should not inherit the same evidential threshold automatically.
Second, name the people. Who places orders, who needs delivery information and who approves spending? In the example, these could be different roles. A generated persona called “purchasing manager” might leave these distinctions unspecified. Record the intended population and the characteristics genuinely known about it, including how that knowledge was obtained.
Third, describe the comparison. Ask the supplier to test predictions against relevant human responses reserved for evaluation. Those answers should not have been supplied to the model to generate the predictions or used to tune the procedure being evaluated. Ask what is known about possible overlap with earlier training data. If it cannot be established, retain that uncertainty in the report.
Fourth, agree what would count as an adequate result before seeing it. For this choice, I would want to know whether the method selects the same preferred feature and whether important user groups lead to different choices. The research design should assess uncertainty and the amount of evidence available for each group. Overall agreement can coexist with an unresolved decision for a smaller group.
Fifth, record the permitted consequence. A result might justify using the simulation to prepare interviews while leaving feature selection dependent on customer research. A stronger local comparison might support a narrowly specified use. Record the model version, instructions, input data and evaluation date so that later changes prompt a decision about renewed evaluation.
Keep the counting units visible
Imagine that 700 of the thousand generated responses favour delivery updates. In this fictional calculation, 70% describes generated responses under the stated procedure. Writing “70% of customers prefer delivery updates” changes the population being described. A customer claim requires evidence of how the procedure estimates preferences within that population, together with the uncertainty of the estimate.
Adding another thousand generated responses may help investigate the procedure’s variability. It does not add a thousand independently observed customers. Likewise, a very narrow interval around a generated proportion would leave uncertainty about its relationship to actual customer preferences. I would require the analyst to explain both questions separately.
Even a well-conducted human preference survey leaves a further distinction: which option people say they prefer, and which one they use when it is available. For our fictional decision, a limited prototype test could investigate use before a larger commitment. That is a proposed next step whose suitability depends on the product and the decision.
Preserve the origin when presenting the result
The 2025 ICC/ESOMAR Code requires disclosure of synthetic data and human oversight to clients; its publication provisions also require methodological transparency. This is a professional code. Disclosure alone establishes no predictive validity. [3]
I would attach the evidence brief to any slide containing the result. A colleague should be able to identify the generated component, the human comparison, the unresolved gaps and the decision authorised. The label should survive when the slide is copied into another presentation.
For the next product meeting, the immediate question is therefore concrete: what would we do differently because of these simulated answers, and what observation supports that change? Where the answer remains open, the simulation can supply questions for the next investigation. The development budget can then follow what that investigation actually establishes.
Sources and further reading
- Argyle et al. (2023)Political Analysis
Political responses.
- Bisbee et al. (2024)Political Analysis
Distributions and relationships.
- ICC/ESOMAR Code (2025)ICC/ESOMAR
Articles 7 and 9.
The starting point for this reflection
Perspective and interests
This article was developed with AI assistance. The company, portal and numerical example are fictional. The evidence brief is an original organisational proposal without scientific validation.
I am the founder and CEO of aiomics. This article reports no customer results and evaluates no current supplier. Sources were checked on 3 October 2026.



