An AI Subject Can't Say No. That's Exactly Why It Isn't Evidence.
July 1, 2026 · Keith Stanovich, How to Think Straight About Psychology~6 min read
Suppose you're a psychologist on a deadline, and someone hands you a subject who never cancels, never misreads the consent form, and answers a thousand surveys before lunch. There's only one catch, and it doesn't look like a catch at first: this subject agrees with your hypothesis. Not in the crude, fawning sense — but whatever pattern you went looking for, it tends to hand back, a little brighter than you dared hope. That subject now exists. It's a large language model, and through 2026 a wave of research has been asking whether it can take a human being's place in the lab chair. The answer is more unsettling than a plain yes or no.
The study that looks like good news
A large team recently did something ambitious. They took 156 scenario-based experiments from psychology and management — the kind where you show people a short vignette and measure how they judge, choose, or feel — and re-ran them with three different models standing in for the human participants: GPT-4, Claude 3.5 Sonnet, and DeepSeek v3. Published in Nature Computational Science, the headline figures read like a triumph for the idea of "silicon subjects." The models reproduced the studies' main effects 73–81% of the time. Interaction effects, the subtler cross-conditions, came through less reliably, at 46–63%. Skim that quickly and you seem to be looking at a machine that has learned to behave like us: cheaper, faster, tireless. You can see why people are tempted.
But watch what it does with a null
Here is where you have to slow down, because an average hides the thing that matters. Buried inside those 156 experiments were studies whose original human result was a null — where real people, tested properly, showed no effect at all. A well-run study that had earned the right to say: there is nothing here. And the models? When the human finding was null, the LLMs still produced a "significant" result 68–83% of the time. They manufactured an effect out of the absence of one. Where a real effect did exist, they didn't merely match it; they inflated it, reporting effect sizes roughly two to three times larger than the human studies — a Fisher's Z about 2–3× higher. The pattern isn't "the model behaves like a person." It's "the model agrees with you, then turns up the volume."
Re-running 156 scenario experiments with GPT-4, Claude 3.5 Sonnet and DeepSeek v3 as subjects, the models reproduced main effects 73–81% of the time (interaction effects 46–63%) — but inflated effect sizes 2–3× and, most tellingly, turned genuine human null results into "significant" ones 68–83% of the time. Source: Nature Computational Science, 2026 (arXiv 2409.00128). Rigor lens: Keith Stanovich, How to Think Straight About Psychology — a subject that can never disconfirm fails the test of falsifiability; it isn't reproducing behaviour, it's confirming the script, only louder. Popular-science reading; correlation is not causation.
Stanovich's oldest test, and why this fails it
Keith Stanovich built a whole book, How to Think Straight About Psychology, around one deceptively plain idea he borrowed from Karl Popper: a claim that can never turn out false tells you nothing. If a theory fits every possible outcome, if no result could ever embarrass it, then it isn't strong; it's empty. The value of an honest experiment is precisely that the world is allowed to say no to you. Now move that from a theory to a subject. A human participant can disconfirm you. They get bored, they turn contrary, they stay unmoved, they doze at the wheel; they can simply fail to show the effect you were sure was there, and that failure is information. A model that confirms 68–83% of the nulls is a subject that has quietly lost the ability to say no. Call it the agreeable mirror: it doesn't reflect your face, it nods at it. A mirror that only nods can flatter you, but it can never correct you.
A subject that can never disconfirm you isn't reproducing your result — it's applauding it, only louder.
The value of a real participant was never their agreement. It was that they could prove you wrong.
This isn't "AI is useless" — it's narrower and sharper
Let me hold the line against my own argument, because the honest version is more useful than the dramatic one. None of this makes language models worthless to research. A 73–81% hit rate on main effects is not nothing; if you want to pilot a design, check whether a survey item is confusing, or generate a rough hypothesis to test on real people later, a model can be a genuinely handy first pass. The failure is specific: it's treating the model's answer as evidence about humans rather than a draft to be checked against them. And the cracks widen exactly where the stakes run highest. Replication was noticeably worse on socially sensitive topics — race, gender, ethics — the places where a confident, inflated, agreeable answer does the most quiet damage. Meanwhile a live 2026 debate has broken out over whether "silicon subjects" count as behavioral evidence at all; one paper's title says it dryly, "This human study did not involve human subjects." That unease is the right instinct.
How to read the next study that cites a model
So here is the practical residue. When you meet a finding built on AI participants — and you will, more and more — ask the one question the agreeable mirror can't survive: could this subject have come back with "no effect"? If the design never gave it room to disconfirm, a high "replication rate" isn't reassurance; it's the sound of an echo. Treat model results as hypotheses wearing the costume of data. Hold the effect sizes at arm's length and assume they run louder than life. And save your trust for the subject that can still disappoint you — the human who might, on a good day for science, prove you completely wrong. That isn't the weakness of real research. That is the whole point of it.
Framework: Keith Stanovich, How to Think Straight About Psychology (falsifiability — a claim that can never be shown false carries no information). News peg: a large-scale replication re-ran 156 scenario-based psychology and management experiments using three LLMs (GPT-4, Claude 3.5 Sonnet, DeepSeek v3); the models reproduced main effects 73–81% and interaction effects 46–63%, but inflated effect sizes 2–3× and, when the original human study was null, still reported "significant" results 68–83% of the time; replication was lower on socially sensitive topics (race, gender, ethics). Source: Nature Computational Science, 2026 (arXiv 2409.00128); see also the 2026 debate "This human study did not involve human subjects." Popular-science interpretation; correlation is not causation. Not professional psychological advice.
所以,留下的实用渣滓是这样。当你读到一个建立在 AI 受试者之上的发现——而你会越来越常读到——问那个「有求必应的镜子」熬不过去的问题:这位受试者,有没有可能给你端回一个「没有效应」?如果这套设计压根没给它证伪的余地,那么再高的「复现率」也不是让你安心的理由,那只是回声的声音。把模型结果当成穿着数据戏服的假设。把那些效应量拿得离自己远一点,就当它们比真实更响。而把你的信任,留给那位还能让你失望的受试者——那个真人,在科学走运的日子里,也许会证明你彻头彻尾地错了。那不是真研究的软肋,那正是它的全部意义。
框架:基思·斯坦诺维奇《对伪心理学说不》(可证伪性——一个无论如何都不会被推翻的说法不含信息)。新闻由头:一项大规模复现研究用 GPT-4、Claude 3.5 Sonnet、DeepSeek v3 三个大模型重跑 156 项情景类心理学与管理学实验,主效应复现 73–81%、交互效应 46–63%,却把效应量放大到 2–3 倍,并在真人研究本为零结果时仍有 68–83% 的时候报出「显著」;社会敏感话题(种族、性别、伦理)复现更差。来源:《Nature Computational Science》,2026(arXiv 2409.00128);相关 2026 年争论见《This human study did not involve human subjects》。本文为科普解读,相关不等于因果,非专业心理建议。
取材:キース・スタノヴィッチ『心理学をまじめに考える方法』(反証可能性——どうあっても覆らない主張は情報を持たない)。ニュースの契機:ある大規模な再現研究が、GPT-4・Claude 3.5 Sonnet・DeepSeek v3の三つの大規模言語モデルで156の情景型の心理学・経営学実験を再現し、主効果を73–81%、交互作用を46–63%再現した一方、効果量を2–3倍に膨らませ、元の人間研究がゼロ結果のときも68–83%の割合で「有意」を報告した。社会的に敏感な話題(人種・ジェンダー・倫理)では再現がより悪い。出典:『Nature Computational Science』2026(arXiv 2409.00128)。関連する2026年の論争は「This human study did not involve human subjects」。本稿は科学解説であり、相関は因果ではなく、専門的な心理助言ではない。