Back
Data quality check

Is your survey data high quality?

Check your data like how social scientists check data quality before analysis. Applicable for both human panel data and simulated data.

Send a datasetHow the check works
High-quality data
Your dataset

Why check internal quality of survey data

And how it applies to both human and simulated respondents.

Pew Research Center found that 4–7% of interviews in online opt-in panels came from 'bogus' respondents — and that traditional quality checks missed most of them: 84% passed the survey's trap question and 87% passed a speeding check. In 2026, Pew's methodologists warned that AI now makes it cheap to fake a respondent at scale in panels that claim to be fully human.

Artificial Societies use large networks of AI to model important groups of people. We're social and behavioral scientists, so we hold our simulations to the same standard as how we would check a human dataset before conducting social science research. A human panel can be contaminated with low-attention respondents or bot-farm participants. A carefully built simulation can be more internally coherent. The only way to tell is to look at the raw, respondent-level data.

What the check looks for

  • Coherent respondents that don't hallucinate and contradict themselves
  • Realistic diversity of opinions by gender, generation and geography
  • Opinions that correlate into themes, the way real attitudes do
  • Open text with depth and nuance, written by many different people

What it doesn't do

  • It doesn't check external accuracy of the findings
  • It is not domain-specific, and focuses on universal internal quality metrics
  • It doesn't tell you whether the data is current, or whether the questions were well designed
  • It doesn't tell you whether the dataset is human or synthetic, just its quality

What your data gets compared against

Metrics social scientists use to check for data quality.

High-quality data has a recognisable shape. Answers spread across the options, differ from question to question, correlate into themes, and vary between different people. Good data is messy with patterns. Low-quality data — rushed, straightlined, automated or generated — can sometimes be too inconsistent, too uniform, or conversely too templated.

AwarenessPurchase intentMotivationRepurchase intentPerceived valueBarriersComfort level
High-quality dataVaried, spread, untidy
Mixed signalsTwo questions spike
Low quality / unlikely humanOne shape, repeated
Share choosing each option
lowerhigher
Illustrative

Too uniform

Individuals barely differ from each other, and no one has a consistent belief across their answers.

The high-quality band

Messy but with pattern: clear themes, clear outliers.

Too templated

Answers collapse into a few repeated shapes, loosing the noise and nuance that make us human.

The four measures of quality

Based on the checks social scientists run before trusting a dataset.

01

Respondent coherence

How often do respondents contradict themselves?

Needs: Questions that are related
02

Subgroup diversity

Do opinions differ by gender, generation, and geography realistically?

Needs: Demographic columns
03

Opinion correlation

Do opinions form thematic structures organically?

Needs: Several rating-scale questions
04

Open text richness

Does the free text have depth and nuance?

Needs: At least one free-text column

What comes back

A short report you can circulate.

A score and a bandHigh quality, mixed signals, low quality / unlikely human, or not assessable.

CoverageHow much of the check could run on your file. A file with no free text can't be scored on open text richness, so we report it as skipped to help contextualise the score.

FindingsOne per measure explaining how to interpret the metric, or reason why the check was skipped.

CaveatsWhat we couldn't measure, and a note that this checks internal quality, not external quality or source of data.

Quality reportIllustrative
61
Mixed signalsout of 100
Coverage3 of 4 checks ran
Respondent coherence11 respondents (2.1%) contradict themselves across related questions.
Subgroup diversityAge and region shift opinions realistically compared to high-quality samples.
Opinion correlationAnswers are correlated into two thematic clusters in realistic ways.
Open text richnessSkipped — the file has no free-text column to read.

Caveat: This report checks the dataset's internal quality, not its accuracy or where the data came from.

What to send

The more of your file arrives intact, the more we can measure.

Raw responses, one row per person

One column per question. We are not able to measure these metrics based on crosstab or topline summaries.

Include the extra columns

Free text unlocks open text richness checks, and demographics unlock subgroup diversity checks.

Strip direct identifiers

Remember to remove names, emails, phone numbers, addresses and panellist IDs first.

Send a dataset for a quality check

Get results within two working days.

Artificial Societies is SOC 2 certified and has strict GDPR compliance. We use your file to run this check and send you the report. Your file will never be shared outside Artificial Societies, and can be deleted on request. See our privacy policy and DPA.

Questions

These are the same quality checks we hold our own simulations to. How we build and evaluate them is written up in our method and evaluation.