Send us a spreadsheet where each row is a person and each column is a question. We measure its quality on four things — coherence, subgroups, correlation and open text — and tell you where it holds up and where it doesn't. Human panel data and synthetic data can each be high or low quality, so the check reads them the same way, on any topic, without ever needing to know whether your numbers are correct.
Rows are questions, columns are answer options, and the darker a cell the more people chose it. We read the shape your answers make, and compare it with the shapes high-quality datasets make.
Poor data isn't only a synthetic-data problem. Plenty of files collected from real people don't capture them well either.
Pew Research Center found that 4–7% of interviews in online opt-in panels came from bogus respondents — and that the usual defences barely caught them: 84% of those respondents passed the survey's trap question and 87% passed a speeding check (Assessing the Risks to Online Polls From Bogus Respondents, 2020). Six years on, Pew's methodologists were still warning that AI makes it cheap to fake a respondent at scale (Do AI and bogus respondents threaten polling's future?, 2026).
So the useful question isn't "is this synthetic?" — it's is this high quality? A panel that promises human respondents can still hand you straightliners, speeders and bots; a carefully built simulation can capture human structure better than a badly run field. Either source can be high quality or low quality, and the only place to settle it is the raw, row-level file.
The benchmark is high-quality data — files where real people answered, and answered properly.
People answering properly leave a particular kind of fingerprint on a dataset. Answers spread across the options, vary from question to question, hang together in themes, and differ between people who have different lives. Good data is untidy in specific, predictable ways. Weak data — generated, rushed, straightlined or automated — is usually too clean, too uniform, or too stereotyped, and it shows up as a different shape.
Rows barely differ from one another. Nobody in the file has a settled outlook you can trace across their answers — where straightliners and inattentive fieldwork land.
Structured but far from tidy — themes you can see, plenty of exceptions. High-quality files live in this middle, whether a panel or a simulation produced them.
Answers collapse into a handful of repeated shapes, cleaner and more stereotyped than any real crowd — where generators and survey bots land.
None of this depends on knowing the right answer. Because the check reads how rows behave rather than what the toplines say, tuning your percentages to look plausible doesn't move it.
Each one is a plain-English question about your file, and each is a way a dataset can turn out to be low quality. Every finding in the report names the number behind it.
Does anyone give two answers that can't both be true?
Needs: Questions that overlap in meaningDo age, region and role predict opinions by a human amount — not zero, not everything?
Needs: Demographic columnsDo opinions hang together in themes, or is everything unrelated — or all one thing?
Needs: Several rating-scale questionsDoes the free text read like many different people wrote it?
Needs: At least one free-text columnA short report you can hand to a colleague or a supplier, written to be read rather than decoded.
A score and a bandHigh quality, mixed signals, low quality or unlikely human, or not assessable.
CoverageHow much of the check could actually run on your file. This is the honesty dial: a file with no free text can't be checked on that, and the report says so rather than quietly scoring it.
FindingsOne per check, in plain English, each naming the specific number behind it. Skipped checks come back with a reason, not a blank.
CaveatsWhat the check couldn't see, and the standing limit that it reads quality, not provenance.
Caveat: a carefully built simulation can pass these tests, and a carelessly run panel can fail them. This report describes patterns, not provenance or intent.
The more of your file survives intact, the more of the check can run.
One column per question. A crosstab or topline summary has no rows to read, so there's nothing for the check to work with.
Don't clean, impute or drop rows before sending. Tidying is exactly what hides the contradictions the check looks for.
Free text unlocks the open-text check and demographics unlock the subgroup check. Missing ones are reported as skipped, not scored.
The check reads answers, not people. Remove names, emails, phone numbers, addresses and panellist IDs before you send.
If answers are numeric codes, tell us what they mean in the notes field — or send a labelled export instead.
CSV, TSV, Excel, JSON, SPSS (.sav) or Stata (.dta), up to 25MB. Bigger than that, or a format not listed — email us and we'll sort it out.
Human panel, synthetic, or your own fieldwork — drop in a file and we'll do the rest. Free, and no obligation to talk to anyone: a researcher on our team runs each check by hand and emails the report back, usually within two working days.
SOC 2 certified and GDPR compliant. We use your file to run this check and send you the report — never to train models, never shared outside Artificial Societies, and deleted on request. See our privacy policy and DPA.
These are the same quality tests we hold our own simulations to. How we build and evaluate them is written up in our method and evaluation.