How well can AI-generated synthetic research data replicate the responses of human participants? This article benchmarks synthetic survey generation against a survey of 400 Silicon Valley coders and developers by asking first for leading consumer AI platforms to produce full synthetic datasets through a naturalistic chat protocol, then testing the three newest frontier models (Claude Fable 5, GPT-5.6 Sol, and Kimi K3). Our findings reveal that while the platforms produced technically plausible results that lean more towards replicability and harmonization than typically assumed, none of the LLMs consistently captured the complete set of counterintuitive findings of the human survey. This pattern varied across different prompt formations and generation methods: deviations in synthetic data grouped together for some model comparisons, while others diverged substantially. While leading LLMs are increasingly being used to scale, replicate and replace human survey responses in innovation and science policy research, the tested models still differed substantially from the human benchmark. Agreement among synthetic samples, whether across models or across runs, therefore does not establish agreement with human populations, showing here that replicability and validity decouple. We propose that synthetic respondents could be used before fieldwork, as an instrument for identifying model-generated expectations about research populations, and we connect this use to emerging calibration methods for settings where a human anchor exists.
Download the full text (PDF)
Final version, published as a White Paper by the Centre for Global Sustainability, University of Oslo (2026), under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 licence that covers its arXiv posting. Also on arXiv: https://doi.org/10.48550/arXiv.2603.00059 and SSRN: https://doi.org/10.2139/ssrn.6210099