Can you trust synthetic users? The BVE Study (Behavioral Validity Evaluation) answers that question with evidence — 11,000+ synthetic interviews across six commercial LLMs and three countries, validated against real consumer transaction data.

The field has been missing large-scale, externally validated evidence that distinguishes which vendor claims hold and which do not. The core message is not “trust them” or “don’t trust them.” A synthetic user is a behavioral instrument — reliable when calibrated for the right question, misleading when it isn’t.

Three findings reshape how the field should think about synthetic populations:

Stimulus design matters 23x more than model choice. 76% of observed variance comes from how you ask, not which LLM you use.

Synthetic agents recover identity-level attributes 2.6x better than operational behaviors. Positioning and attitudinal segmentation: strong. SKU-level forecasting and switching prediction: weak.

More context does not mean more accuracy. Adding demographic and OCEAN layers to rich behavioral signal can degrade predictive fidelity.

The question for practitioners is no longer “should we use synthetic data?” but “for which question, with which calibration, and against which validation?”

Read the complete Paper here.