Stochastic Parrots or Singing in Harmony? Testing Leading AI LLMs’ Ability to Replicate a Human Survey with Synthetic Data

Jason Miklian, Kristian Hoelscher, John E. Katsos

White Paper, Centre for Global Sustainability, University of Oslo, 2026

This page and its PDF are the white paper of 1 October 2026 (Zenodo DOI 10.5281/zenodo.23087156). An earlier version, titled “Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data” and reporting 420 respondents, is on arXiv (2603.00059, February 2026) and SSRN · Synthetic dataset (Hugging Face) · Replication package (zip)

Free PDF Publisher version → Cite

How well can AI-generated synthetic research data replicate the responses of human participants? This article benchmarks synthetic survey generation against a survey of 400 Silicon Valley coders and developers by asking first for leading consumer AI platforms to produce full synthetic datasets through a naturalistic chat protocol, then testing the three newest frontier models (Claude Fable 5, GPT-5.6 Sol, and Kimi K3). Our findings reveal that while the platforms produced technically plausible results that lean more towards replicability and harmonization than typically assumed, none of the LLMs consistently captured the complete set of counterintuitive findings of the human survey. This pattern varied across different prompt formations and generation methods: deviations in synthetic data grouped together for some model comparisons, while others diverged substantially. While leading LLMs are increasingly being used to scale, replicate and replace human survey responses in innovation and science policy research, the tested models still differed substantially from the human benchmark. Agreement among synthetic samples, whether across models or across runs, therefore does not establish agreement with human populations, showing here that replicability and validity decouple. We propose that synthetic respondents could be used before fieldwork, as an instrument for identifying model-generated expectations about research populations, and we connect this use to emerging calibration methods for settings where a human anchor exists.
Download the full text (PDF)

Final version, published as a White Paper by the Centre for Global Sustainability, University of Oslo (2026), under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 licence that covers its arXiv posting. Also on arXiv: https://doi.org/10.48550/arXiv.2603.00059 and SSRN: https://doi.org/10.2139/ssrn.6210099

Key Messages

  • The benchmark is a survey of Silicon Valley coders and developers with 400 complete responses. Study 1 asked consumer versions of ChatGPT, Claude, Gemini and DeepSeek, plus Claude Cowork, for full synthetic datasets through a naturalistic chat protocol between October 2025 and January 2026; Study 2 tested Claude Fable 5, GPT-5.6 Sol and Kimi K3 in July 2026 under both the chat protocol and respondent-level generation.
  • No model and no generation method consistently recovered the complete set of the human survey’s counterintuitive findings, although some individual rates came close. All three comparable Study 1 platforms put regret about a product’s social impact below the human 50.8%, at 9.9% to 39.5%.
  • The generation method changed the answers. Moving from the chat protocol to respondent-level generation pushed compliance with a freedom-restricting directive to 99.2% for Fable 5 and 99.9% for Kimi K3 and cut it to 14.6% for GPT-5.6, against a human 63.5%; regret went from undershooting in all seven complete chat runs (19.0% to 30.2%) to overshooting in all nine respondent-level runs (61% to 82%).
  • Replicability and validity decouple. Models agreed with themselves across runs more than with each other (mean Jensen-Shannon distance 0.039 within models against 0.252 between models under respondent-level generation), and neither kind of agreement established agreement with the human respondents.
  • The authors propose using synthetic respondents before fieldwork, to surface the expectations a model holds about a population, and set out reporting standards: generation method, multiple independent runs, prompt-sensitivity ranges, exact model versions with dates and settings, and contamination checks against any human benchmark claimed as novel.

Data and Replication

Available responses, prompts, run logs, coding and analysis scripts, tables, figure-generation code, and the provenance probe log are supplied in the accompanying replication package. Download: Parrots replication package v2.1 (zip, Google Drive). The 10,592 synthetic respondents, generation materials and codebook are also published as a dataset on Hugging Face: miklia/llm-synthetic-survey-respondents-stochastic-parrots.

Research Topics

LLMs synthetic data AI research methods survey methodology silicon sampling computational social science

Citation

Jason Miklian, Kristian Hoelscher, John E. Katsos. “Stochastic Parrots or Singing in Harmony? Testing Leading AI LLMs’ Ability to Replicate a Human Survey with Synthetic Data.” White Paper, Centre for Global Sustainability, University of Oslo, 2026.

Related Research Areas

BibTeX Citation
@article{miklian2026_stochastic_parrots_or_singing,
  title = {Stochastic Parrots or Singing in Harmony? Testing Leading AI LLMs’ Ability to Replicate a Human Survey with Synthetic Data},
  author = {Miklian, Jason and Hoelscher, Kristian and Katsos, John E.},
  journal = {White Paper, Centre for Global Sustainability, University of Oslo},
  year = {2026},
  doi = {10.5281/zenodo.23087156},
  eprint = {2603.00059},
  archiveprefix = {arXiv},
  url = {https://miklian.org/papers/stochastic-parrots-or-singing-in-harmony-testing-five-leading-llms-for-their/}
}