vlog
← 返回全部文章

Psychology

An AI Subject Can't Say No. That's Exactly Why It Isn't Evidence.

July 1, 2026 · Keith Stanovich, How to Think Straight About Psychology~6 min read

Suppose you're a psychologist on a deadline, and someone hands you a subject who never cancels, never misreads the consent form, and answers a thousand surveys before lunch. There's only one catch, and it doesn't look like a catch at first: this subject agrees with your hypothesis. Not in the crude, fawning sense — but whatever pattern you went looking for, it tends to hand back, a little brighter than you dared hope. That subject now exists. It's a large language model, and through 2026 a wave of research has been asking whether it can take a human being's place in the lab chair. The answer is more unsettling than a plain yes or no.

The study that looks like good news

A large team recently did something ambitious. They took 156 scenario-based experiments from psychology and management — the kind where you show people a short vignette and measure how they judge, choose, or feel — and re-ran them with three different models standing in for the human participants: GPT-4, Claude 3.5 Sonnet, and DeepSeek v3. Published in Nature Computational Science, the headline figures read like a triumph for the idea of "silicon subjects." The models reproduced the studies' main effects 73–81% of the time. Interaction effects, the subtler cross-conditions, came through less reliably, at 46–63%. Skim that quickly and you seem to be looking at a machine that has learned to behave like us: cheaper, faster, tireless. You can see why people are tempted.

But watch what it does with a null

Here is where you have to slow down, because an average hides the thing that matters. Buried inside those 156 experiments were studies whose original human result was a null — where real people, tested properly, showed no effect at all. A well-run study that had earned the right to say: there is nothing here. And the models? When the human finding was null, the LLMs still produced a "significant" result 68–83% of the time. They manufactured an effect out of the absence of one. Where a real effect did exist, they didn't merely match it; they inflated it, reporting effect sizes roughly two to three times larger than the human studies — a Fisher's Z about 2–3× higher. The pattern isn't "the model behaves like a person." It's "the model agrees with you, then turns up the volume."

AN AGREEABLE MIRROR CAN'T SAY NO — SO IT ISN'T EVIDENCE 156 experimentsGPT-4 · Claude · DeepSeekstand in as subjects WHAT THE MIRROR REPORTS main effects reproduced73–81% of the time effect sizes inflated2–3× louder than humans THE TELLhuman nulls → "significant"68–83% of the time an agreeable mirrorthat can never say no→ carries no information a real subject can say "no"that's the whole value of testing one 156 scenario experiments · Nature Computational Science, 2026.A benchmark of imitation, not evidence of human behaviour. reported flow what a real subject offers
Re-running 156 scenario experiments with GPT-4, Claude 3.5 Sonnet and DeepSeek v3 as subjects, the models reproduced main effects 73–81% of the time (interaction effects 46–63%) — but inflated effect sizes 2–3× and, most tellingly, turned genuine human null results into "significant" ones 68–83% of the time. Source: Nature Computational Science, 2026 (arXiv 2409.00128). Rigor lens: Keith Stanovich, How to Think Straight About Psychology — a subject that can never disconfirm fails the test of falsifiability; it isn't reproducing behaviour, it's confirming the script, only louder. Popular-science reading; correlation is not causation.

Stanovich's oldest test, and why this fails it

Keith Stanovich built a whole book, How to Think Straight About Psychology, around one deceptively plain idea he borrowed from Karl Popper: a claim that can never turn out false tells you nothing. If a theory fits every possible outcome, if no result could ever embarrass it, then it isn't strong; it's empty. The value of an honest experiment is precisely that the world is allowed to say no to you. Now move that from a theory to a subject. A human participant can disconfirm you. They get bored, they turn contrary, they stay unmoved, they doze at the wheel; they can simply fail to show the effect you were sure was there, and that failure is information. A model that confirms 68–83% of the nulls is a subject that has quietly lost the ability to say no. Call it the agreeable mirror: it doesn't reflect your face, it nods at it. A mirror that only nods can flatter you, but it can never correct you.

A subject that can never disconfirm you isn't reproducing your result — it's applauding it, only louder.

The value of a real participant was never their agreement. It was that they could prove you wrong.

This isn't "AI is useless" — it's narrower and sharper

Let me hold the line against my own argument, because the honest version is more useful than the dramatic one. None of this makes language models worthless to research. A 73–81% hit rate on main effects is not nothing; if you want to pilot a design, check whether a survey item is confusing, or generate a rough hypothesis to test on real people later, a model can be a genuinely handy first pass. The failure is specific: it's treating the model's answer as evidence about humans rather than a draft to be checked against them. And the cracks widen exactly where the stakes run highest. Replication was noticeably worse on socially sensitive topics — race, gender, ethics — the places where a confident, inflated, agreeable answer does the most quiet damage. Meanwhile a live 2026 debate has broken out over whether "silicon subjects" count as behavioral evidence at all; one paper's title says it dryly, "This human study did not involve human subjects." That unease is the right instinct.

How to read the next study that cites a model

So here is the practical residue. When you meet a finding built on AI participants — and you will, more and more — ask the one question the agreeable mirror can't survive: could this subject have come back with "no effect"? If the design never gave it room to disconfirm, a high "replication rate" isn't reassurance; it's the sound of an echo. Treat model results as hypotheses wearing the costume of data. Hold the effect sizes at arm's length and assume they run louder than life. And save your trust for the subject that can still disappoint you — the human who might, on a good day for science, prove you completely wrong. That isn't the weakness of real research. That is the whole point of it.

Framework: Keith Stanovich, How to Think Straight About Psychology (falsifiability — a claim that can never be shown false carries no information). News peg: a large-scale replication re-ran 156 scenario-based psychology and management experiments using three LLMs (GPT-4, Claude 3.5 Sonnet, DeepSeek v3); the models reproduced main effects 73–81% and interaction effects 46–63%, but inflated effect sizes 2–3× and, when the original human study was null, still reported "significant" results 68–83% of the time; replication was lower on socially sensitive topics (race, gender, ethics). Source: Nature Computational Science, 2026 (arXiv 2409.00128); see also the 2026 debate "This human study did not involve human subjects." Popular-science interpretation; correlation is not causation. Not professional psychological advice.

心理

AI 受试者永远说不出「不」——这正是它当不了证据的原因

2026年7月1日 · 基思·斯坦诺维奇《对伪心理学说不》约 5 分钟

假设你是个赶着交差的心理学研究者,有人给你送来一位受试者:从不爽约,从不看错知情同意书,一个上午能答完上千份问卷。只有一个麻烦,而且乍看根本不像麻烦——这位受试者,总是同意你的假设。不是那种露骨的迎合,而是你去找什么规律,它多半就给你端回什么,还比你原本敢期望的更亮眼几分。这样的受试者如今真的存在了。它是一个大语言模型,而整个 2026 年,一波接一波的研究都在追问:它能不能替一个真人,坐进实验室的那把椅子。答案,比简单的能或不能更让人不安。

看起来像好消息的那项研究

最近,一支庞大的团队做了件雄心勃勃的事。他们把心理学和管理学里的 156 项情景实验——就是给人看一段小情境,再量他怎么判断、怎么选、怎么感受的那类——拿三个不同的模型替真人重跑了一遍:GPT-4、Claude 3.5 Sonnet 和 DeepSeek v3。这项研究发在《Nature Computational Science》上,几个头条数字读起来简直是「硅基受试者」这个想法的凯歌。模型把这些研究的主效应复现了 73–81%。交互效应——那些更微妙的交叉条件——就没那么可靠了,只有 46–63%。快快扫一眼,这像是一台已经学会像我们一样反应的机器:更便宜、更快、不知疲倦。你不难理解,为什么有人动心。

可是,看看它拿一个零结果怎么办

就在这儿你得慢下来,因为平均数会盖住真正要紧的东西。那 156 项实验里,埋着一些原本真人结果是零的研究——真人被规规矩矩地测过,什么效应也没有。空的。一项做得扎实的研究,挣来了说「这里什么都没有」的资格。那模型呢?当真人的结论是零时,这些大模型仍有 68–83% 的时候,报出一个「显著」的结果。它们从「没有」里,硬造出了一个「有」。而在真有效应的地方,它们不只是对上了,还把它吹大:报出的效应量大约是真人研究的两到三倍——Fisher's Z 高出约 2–3 倍。这个模式不是「模型表现得像个人」,而是「模型先同意你,再把音量拧大」。

有求必应的镜子说不出「不」——所以它不是证据 156 项实验GPT-4 · Claude · DeepSeek被当作受试者 镜子报告了什么 主效应被复现73–81% 的比例 效应量被夸大比真人高 2–3 倍 破绽在这里真人的零结果 →「显著」68–83% 的时候 有求必应的镜子永远说不出「不」→ 不含任何信息 真受试者能说「不」这才是测一个人的全部价值 156 项情景实验 · Nature Computational Science,2026。这是模仿的基准,不是人类行为的证据。 被报告的流向 真受试者能给的
用 GPT-4、Claude 3.5 Sonnet、DeepSeek v3 当受试者重跑 156 项情景实验,模型把主效应复现了 73–81%(交互效应 46–63%)——却把效应量放大到 2–3 倍;最露馅的是,当真人研究本是零结果时,模型仍有 68–83% 的时候报出「显著」。来源:《Nature Computational Science》,2026(arXiv 2409.00128)。严谨度视角:基思·斯坦诺维奇《对伪心理学说不》——一个永远无法被证伪的受试者过不了可证伪性这一关;它不是在复现行为,而是把你写好的剧本再确认一遍,只是更响。科普解读;相关不等于因果。

斯坦诺维奇那道最老的考题,它没过

基思·斯坦诺维奇写了一整本《对伪心理学说不》,就绕着一个看似朴素、借自卡尔·波普尔的念头:一个永远不可能被证明为假的说法,什么也没告诉你。如果一套理论跟任何可能的结果都相容,如果没有任何结果能让它难堪,那它不是强,而是空。一个诚实实验的价值恰恰在于:世界被允许对你说「不」。现在把这条从理论挪到受试者身上。一个真人能证伪你。他会走神、会拧着来、会无动于衷、会开着车打瞌睡;他可能就是没表现出你笃定存在的那个效应,而这个「没有」本身,就是信息。一个把 68–83% 的零结果都确认成「显著」的模型,是一位悄悄丧失了说「不」这个能力的受试者。就叫它「有求必应的镜子」吧:它不照你的脸,它冲你点头。而一面只会点头的镜子,能奉承你,却永远纠正不了你。

一个永远无法证伪你的受试者,不是在复现你的结果——它是在给它鼓掌,只是更响。

真受试者的价值,从来不是他同意你。而是他能证明你错了。

这不是「AI 没用」——它的意思更窄,也更锋利

让我按住自己的论点,因为诚实的那个版本,比戏剧化的版本更有用。上面这些,都不等于语言模型对研究毫无价值。主效应 73–81% 的命中率并不是零;你若想给一个设计做预演、检查某道问卷题目是不是让人犯迷糊、或先生成一个粗糙假设留着日后拿真人去验,模型确实能当一次挺顺手的初稿。问题很具体:错在把模型的答案当成关于人类的证据,而不是一份还得拿真人去核对的草稿。而裂缝恰恰在最要紧的地方张得最大。在社会敏感话题上——种族、性别、伦理——复现明显更差,而那正是一个自信、被夸大、又有求必应的答案最会悄悄造成伤害的地方。与此同时,2026 年一场活生生的争论也炸开了:「硅基受试者」到底算不算行为证据?有篇论文的标题说得干脆——「这项人类研究并未涉及人类受试者」。这份不安,是对的直觉。

下次再看到引用模型的研究,你该怎么读

所以,留下的实用渣滓是这样。当你读到一个建立在 AI 受试者之上的发现——而你会越来越常读到——问那个「有求必应的镜子」熬不过去的问题:这位受试者,有没有可能给你端回一个「没有效应」?如果这套设计压根没给它证伪的余地,那么再高的「复现率」也不是让你安心的理由,那只是回声的声音。把模型结果当成穿着数据戏服的假设。把那些效应量拿得离自己远一点,就当它们比真实更响。而把你的信任,留给那位还能让你失望的受试者——那个真人,在科学走运的日子里,也许会证明你彻头彻尾地错了。那不是真研究的软肋,那正是它的全部意义。

框架:基思·斯坦诺维奇《对伪心理学说不》(可证伪性——一个无论如何都不会被推翻的说法不含信息)。新闻由头:一项大规模复现研究用 GPT-4、Claude 3.5 Sonnet、DeepSeek v3 三个大模型重跑 156 项情景类心理学与管理学实验,主效应复现 73–81%、交互效应 46–63%,却把效应量放大到 2–3 倍,并在真人研究本为零结果时仍有 68–83% 的时候报出「显著」;社会敏感话题(种族、性别、伦理)复现更差。来源:《Nature Computational Science》,2026(arXiv 2409.00128);相关 2026 年争论见《This human study did not involve human subjects》。本文为科普解读,相关不等于因果,非专业心理建议。

心理学

「ノー」と言えないAI被験者は、だからこそ証拠にならない

2026年7月1日 · キース・スタノヴィッチ『心理学をまじめに考える方法』約 7 分

締め切りに追われる心理学者だと想像してほしい。誰かが一人の被験者を連れてくる。ドタキャンしない、同意書を読み違えない、昼までに千件のアンケートに答える。難点はひとつだけ、しかも最初は難点に見えない——この被験者は、あなたの仮説に同意する。露骨なへつらいではない。あなたが探しに行ったどんなパターンも、たいてい返してくれる。しかも、望んでいたより少し明るく。そんな被験者が、いまや本当に存在する。大規模言語モデルだ。そして2026年、研究の波は次々と問うている——それは実験室のあの椅子に、人間の代わりに座れるのか、と。答えは、単純なイエスやノーより、ずっと落ち着かない。

良い知らせに見える研究

最近、ある大規模なチームが野心的なことをやった。心理学と経営学から156の情景実験——小さな筋書きを見せて、どう判断し、選び、感じるかを測る、あの種類——を取り、三つの異なるモデルに人間参加者の代わりをさせて再現した。GPT-4、Claude 3.5 Sonnet、そしてDeepSeek v3だ。『Nature Computational Science』に載ったその見出しの数字は、「シリコン被験者」という発想の凱歌のように読める。モデルは各研究の主効果を73–81%の割合で再現した。交互作用——より微妙な条件の掛け合わせ——はそこまで頼りにならず、46–63%だった。ざっと読めば、これは私たちのように振る舞うことを覚えた機械に見える。安くて、速くて、疲れ知らずの。人が惹かれるのも無理はない。

だが、一つのゼロ結果をどう扱うかを見よ

ここで足を止めねばならない。平均は、いちばん大事なものを覆い隠すからだ。あの156の実験の中に、元の人間の結果がゼロだった研究が埋もれていた——本物の人間がきちんと測られて、何の効果も出なかった。空だ。よく設計された研究が、「ここには何もない」と言う資格を勝ち取った、ということだ。ではモデルは?人間の結論がゼロのとき、大規模言語モデルは68–83%の割合で、なお「有意」な結果を吐き出した。「ない」から「ある」をこしらえたのだ。そして本物の効果があった場面では、ただ一致させるだけでなく、膨らませた。報告された効果量は人間研究のおよそ二倍から三倍——Fisher's Zで約2–3倍高い。この型は「モデルが人のように振る舞う」ではない。「モデルはまず同意し、それから音量を上げる」だ。

「ノー」と言えない鏡は、だから証拠にならない 156の実験GPT-4・Claude・DeepSeekを被験者に 鏡は何を報告したか 主効果を再現73–81% の割合で 効果量を誇張人間の2–3倍に 綻びはここ人間のゼロ結果 →「有意」68–83% の割合で 何にでも頷く鏡「ノー」と言えない→ 情報を持たない 本物の被験者は「ノー」と言えるそれこそ被験者を試す値打ちだ 156の情景実験 · Nature Computational Science、2026。模倣のベンチマークであって、人間行動の証拠ではない。 報告された流れ 本物の被験者が与えるもの
GPT-4・Claude 3.5 Sonnet・DeepSeek v3を被験者に156の情景実験を再現したところ、モデルは主効果を73–81%再現した(交互作用は46–63%)——だが効果量を2–3倍に誇張し、そして何より、元の人間研究がゼロ結果だった場面でも68–83%の割合で「有意」を作り出した。出典:『Nature Computational Science』2026(arXiv 2409.00128)。厳密さの視点:キース・スタノヴィッチ『心理学をまじめに考える方法』——決して反証できない被験者は反証可能性の検査を通らない。行動を再現しているのではなく、書かれた台本をより大きな声で確認しているだけだ。科学解説;相関は因果ではない。

スタノヴィッチの最も古い問い、それに落ちる

キース・スタノヴィッチは『心理学をまじめに考える方法』一冊を、カール・ポパーから借りた一見素朴な考えのまわりに築いた——決して偽と分かりようのない主張は、何も教えてくれない、と。ある理論があらゆる可能な結果と両立するなら、どんな結果もそれを困らせられないなら、それは強いのではなく、空っぽだ。誠実な実験の値打ちは、まさに、世界があなたに「ノー」と言うのを許されている点にある。これを理論ではなく被験者に当てはめよう。人間の参加者は、あなたを反証できる。退屈し、へそを曲げ、心を動かされず、居眠り運転をする。あなたが確信していた効果を出さないことがあり、その「出なかった」こと自体が情報だ。ゼロ結果の68–83%を「有意」と確認してしまうモデルは、「ノー」と言う力をひそかに失った被験者である。これを「何にでも頷く鏡」と呼ぼう。あなたの顔を映すのではなく、あなたに頷く。そして頷くだけの鏡は、世辞は言えても、あなたを正すことは決してできない。

決してあなたを反証できない被験者は、あなたの結果を再現しているのではない——拍手を送っているのだ、ただし、より大きな声で。

本物の参加者の値打ちは、同意してくれることではなかった。あなたが間違っていると証明できることだった。

これは「AIは役立たず」ではない——もっと狭く、もっと鋭い

自分の論を、自分で押さえておこう。誠実な版のほうが、劇的な版より役に立つからだ。以上のどれも、言語モデルが研究に無価値だという意味ではない。主効果の73–81%という的中率はゼロではない。設計の下見をしたい、あるアンケート項目が紛らわしくないか点検したい、あとで本物の人間にぶつける粗い仮説を出したい——そんなとき、モデルは本当に手軽な下書きになりうる。失敗は具体的だ。モデルの答えを、人間について確かめるべき下書きではなく、人間についての証拠として扱うこと、そこが誤りだ。そして裂け目は、賭け金がいちばん高いところで最も広がる。社会的に敏感な話題——人種、ジェンダー、倫理——では再現が目に見えて悪く、そここそ、自信に満ちて誇張された、何にでも頷く答えが、最も静かに害をなす場所だ。同じころ2026年、生きた論争も噴き出した。「シリコン被験者」はそもそも行動の証拠に数えられるのか、と。ある論文の題はそっけなく言う——「この人間研究に人間の被験者は関与していない」。その落ち着かなさは、正しい勘だ。

次にモデルを引く研究を読むとき、どうするか

だから、実用の残りかすはこうだ。AI被験者の上に築かれた発見を読むとき——そしてあなたはこれから、ますます読むことになる——あの「何にでも頷く鏡」が生き延びられない問いを投げよ。この被験者は、「効果なし」と返してくる余地があったのか、と。設計が最初から反証の余地を与えていないなら、高い「再現率」は安心の材料ではない。それはこだまの音だ。モデルの結果は、データの衣装をまとった仮説として扱え。効果量は腕一本ぶん遠ざけて持ち、実物より大きな声だと思っておけ。そして信頼は、まだあなたを失望させられる被験者のために取っておけ——科学に運のいい日には、あなたが完全に間違っていると証明してくれるかもしれない、あの人間のために。それは本物の研究の弱みではない。それこそが、その全目的なのだ。

取材:キース・スタノヴィッチ『心理学をまじめに考える方法』(反証可能性——どうあっても覆らない主張は情報を持たない)。ニュースの契機:ある大規模な再現研究が、GPT-4・Claude 3.5 Sonnet・DeepSeek v3の三つの大規模言語モデルで156の情景型の心理学・経営学実験を再現し、主効果を73–81%、交互作用を46–63%再現した一方、効果量を2–3倍に膨らませ、元の人間研究がゼロ結果のときも68–83%の割合で「有意」を報告した。社会的に敏感な話題(人種・ジェンダー・倫理)では再現がより悪い。出典:『Nature Computational Science』2026(arXiv 2409.00128)。関連する2026年の論争は「This human study did not involve human subjects」。本稿は科学解説であり、相関は因果ではなく、専門的な心理助言ではない。