How Accurate Are AI Personas?

AI persona accuracy is measured along several dimensions — most importantly, how closely simulated opinion distributions match real human opinion distributions, and how consistently individual personas maintain coherent internal belief systems. Artificial Societies reaches 95% of the human self-replication level on distribution accuracy (86% absolute versus a 91% human self-replication ceiling), keeps hallucination under 2%, achieves 89% internal coherence, and produces open-ended responses with 93% of human-level richness. For the full methodology and benchmarks, see our Method & Evaluation.

What Is Opinion Distribution Accuracy?

Opinion distribution accuracy measures how closely the aggregate responses of a simulated audience match the distribution of responses from real humans on the same questions. If 60% of real humans agree with a statement, an accurate simulation should produce a similar proportion. Artificial Societies reaches 95% of the human self-replication level — the rate at which humans themselves replicate their own survey responses (86% absolute accuracy against a 91% human ceiling). Because human survey data contains inherent noise, the self-replication ceiling, not 100%, is the meaningful target.

What Is Persona Internal Coherence?

Persona internal coherence measures whether a single persona maintains a consistent, realistic belief system across many different questions. A coherent persona's responses on politics, economics, social values, and personal priorities should reflect a unified worldview — just as a real person's views are shaped by their underlying values. Artificial Societies achieves 89% coherence (Cronbach's α), within the 60–95% range human panels span and comfortably above the 70% threshold for acceptable quality. This is what enables qualitative interviews, individual-level analysis, and multi-question survey designs — most approaches report distribution accuracy alone and never measure coherence.

What Other Dimensions of Accuracy Matter?

Distribution accuracy and coherence are necessary but not sufficient. Artificial Societies also measures hallucination rate — how often participants contradict themselves — keeping it under 2%, better than the roughly 9% seen in human panels and far below the up-to-35% of biography-prompted LLMs. Open-response quality, benchmarked against 120,000 human social-media posts, reaches 93% of human-level distinct-word richness. Our Method & Evaluation page details how each of these is measured against human panels and simple synthetic personas.

Frequently Asked Questions

How do you measure AI persona accuracy?

Artificial Societies measures accuracy across several dimensions: distribution accuracy (95% of the human self-replication level), internal coherence (89% Cronbach's α), hallucination rate (under 2%), and open-response quality (93% of human-level richness). This is more rigorous than approaches that report distribution accuracy alone. The full methodology is documented on our Method & Evaluation page.

What is the human self-replication ceiling?

The human self-replication ceiling is the rate at which real humans replicate their own survey responses when asked the same questions twice. Human survey data contains inherent noise, so even real humans do not reach 100% self-replication — the ceiling sits around 91%. It represents the meaningful maximum for any audience simulation, which is why Artificial Societies benchmarks against it rather than against a raw 100%.

Related Topics