vlog
← 返回全部文章

Psychology

The AI That Aced 160 Psychology Tests — By Not Reading Them

July 12, 2026 · Dana Cole · Daniel Kahneman, Thinking, Fast and Slow~13 min read

You know the feeling where you've just finished marking a stack of perfectly correct exam papers, and then — in some thought experiment that psychologists love — you imagine asking the student who aced them to simply pick option A on the next question, no matter what. And the student, who is apparently the most capable student you've ever taught, ignores you completely and picks the answer that would score highest on a normal exam. Would you still say they'd understood anything? That's roughly — and in surprisingly precise, published form — what happened to Centaur, the AI model that Nature declared capable of simulating human cognition across 160 behavioural tasks. The declaration was premature, which is quite a lot to get wrong in a flagship journal.

1

Nature Said "It Simulates Human Cognition" — 160 Tasks, Near-Perfect Scores, Worldwide Headlines

In July 2025, a model called Centaur arrived in the pages of Nature with a bold claim attached.

The researchers behind Centaur — led by Marcel Binz and colleagues — had done something genuinely unusual. Instead of fine-tuning a large language model on internet text or task accuracy, they'd trained it specifically on behavioural data from 160 cognitive psychology experiments. The range was real: decision-making under uncertainty, executive control, memory retrieval, risk assessment, learning from feedback. Decades of careful human-subjects research, all the established patterns that distinguish systematic thinkers from impulsive ones, everything that maps onto the processes cognitive psychologists have been documenting for generations.

After fine-tuning on that corpus, Centaur's outputs matched human behavioural patterns with a precision that genuinely impressed people who are hard to impress. It didn't just get the right answers — it got things wrong in the same ways humans do. Same response-time distributions, same biases, same characteristic swerves away from pure rationality that researchers have been cataloguing since the seventies. The paper's language was careful: this model, the authors argued, could serve as a "foundation model of human cognition" — a computational stand-in for how real minds actually process the world.

The headlines were not careful. "AI Thinks Like a Human," several outlets reported. For anyone who had spent years watching the slow, incremental grind of cognitive modelling, a model that could replicate human behaviour across 160 varied tasks felt like a qualitative leap. The claim moved fast. It turned out to have travelled much further than the evidence warranted.

The Centaur claim: Fine-tuning on 160 behavioural psychology experiments produced an AI model that matches human behavioural patterns with high precision across a diverse range of cognitive tasks — suggesting it could function as a computational model of human cognition.
2

Then Researchers Said: "Fine. Now Choose Option A." — And It Didn't.

The test that exposed everything was almost insultingly simple.

Researchers from Zhejiang University, writing in National Science Open in 2026, decided to probe Centaur's claimed understanding with a very straightforward intervention. In the original psychology experiments, participants receive task instructions and then respond accordingly. What if you changed the instructions explicitly? What if you told the model, in plain language, that it must select a specific response — say, option A — regardless of what the task appears to call for?

A human participant, given that instruction, would comply. They might find it strange (I would, honestly, and I'd probably ask why), but they'd follow it. Following an explicit instruction that overrides your default behaviour is one of the clearest markers of what Kahneman calls System 2 operation: deliberate, effortful, instruction-sensitive processing that can override automatic responses. Four- and five-year-olds can do it. It's not a sophisticated cognitive achievement. It's basic instruction comprehension.

Centaur could not do it. When told to choose option A, the model continued selecting whichever response the training distribution indicated was most probable. It wasn't reading the instruction and deciding to ignore it — it was, in the most literal sense, not processing the instruction at all. ScienceDaily's summary captured the finding in a headline that landed hard: "This AI knew the answers but didn't understand the questions."

The Zhejiang team's interpretation was unambiguous. Centaur's remarkable match to human behavioural data was the product of overfitting — the model hadn't learned anything about cognitive processes; it had learned the statistical regularities of a particular training corpus. When the input deviated from those regularities in a way that required actual instruction processing, the whole façade collapsed. "Can Centaur truly simulate human cognition?" the paper asked in its title. The answer it provided: no, not in any meaningful sense.

What Centaur did

When explicitly told "you must choose option A," continued selecting whichever response was statistically most probable given its training distribution. Ignored the instruction entirely.

What genuine cognition would require

Read the instruction. Recognise that the instruction overrides default behaviour. Select option A. This is routine for any human participant — or any system that actually processes what it's told.

3

The System 1 / System 2 Divide: Getting the Answer Right Is Not the Same as Understanding the Question

Kahneman's two-system framework was built precisely to make this distinction legible.

In Thinking, Fast and Slow, Daniel Kahneman describes two modes of cognitive operation that are almost like two different characters sharing a single brain. System 1 is the automatic one — it runs constantly, effortlessly, below conscious awareness. It recognises faces, reads emotional expressions, retrieves answers to familiar questions, matches new inputs to stored patterns. System 1 is fast because it never really processes anything fresh; it's a sophisticated retrieval machine that looks at what's in front of it and returns the most statistically likely response from everything it has encountered before. It's very good at this. Most of the time, that's exactly what you need.

System 2 is the effortful, deliberate one. It's slow, expensive, and prone to wandering off. It kicks in when System 1 gets stuck — when the automatic answer feels off, when the task demands explicit rule-following, when instructions need to be held in working memory and applied step by step. The clearest test for System 2, in Kahneman's framework, is exactly the kind of instruction-following that the Zhejiang team used: given an explicit rule that overrides your default, can you apply it? System 2 is the part of you that can be told "this time, do the opposite" and actually do the opposite. System 1 can't. It only does what it always does.

Centaur is a System 1 machine of extraordinary scale. Trained on 160 experiments' worth of human behavioural data, its System 1 is exquisitely calibrated to human behavioural patterns. But it has no System 2 at all — no mechanism for holding an explicit instruction in working memory and applying it to override the probability distribution it's learnt. When Centaur "matched human cognition" across 160 tasks, it was doing something more limited: it was matching the statistical distribution of human responses in those tasks. That's a real and impressive thing to do. It is not the same thing as simulating cognition.

The distinction matters because human cognition routinely requires both. When you take a multiple-choice test, you're using System 1 (rapid pattern-matching to retrieve relevant knowledge) and System 2 (checking your answers, noticing when something feels right for the wrong reasons, following any special instructions about format or timing). Centaur only ever uses the first half.

ANSWERING ≠ UNDERSTANDING · WHERE CENTAUR LIVES AND WHAT IT CANNOT REACH SYSTEM 1 · FAST · AUTOMATIC pattern-matching · statistical retrieval · no conscious effort Human Intuition fast, automatic pattern recognition Centaur (AI) trained on 160 tasks statistical retrieval only Both live here ↑ Centaur matches human S1 responses with extraordinary precision SYSTEM 2 · SLOW · DELIBERATE instruction-following · rule application · effortful override Human Understanding reads instructions, applies rules can override default response Only humans here ↑ Centaur cannot follow explicit instructions that override training Not there yet dashed = the gap Centaur cannot cross psych domain color System 2 territory (gold) Centaur's unreachable boundary
Conceptual map of System 1 / System 2 territory (Kahneman, Thinking, Fast and Slow) and where Centaur sits. Both human intuition and Centaur operate in System 1 — fast, pattern-based, automatic retrieval. Only humans occupy System 2: deliberate, instruction-following, capable of overriding trained defaults. The dashed arrow marks the boundary Centaur's instruction-ignore failure exposed. Diagram is a conceptual illustration; not a measured cognitive map.
4

The Intelligence Metrics We Use Were Built to Measure System 1

This is the structural problem hiding beneath the Centaur debate: our benchmarks were never designed to test understanding.

Think about what most AI cognitive benchmarks actually measure. A model is given a question and evaluated on whether its answer matches the correct answer (or, in the case of psychology tasks, the typical human answer). The evaluation is outcome-based: right output or wrong output. There's no mechanism in the benchmark to distinguish between "this model understood the question and reasoned its way to the correct answer" and "this model retrieved the statistically most likely answer for this type of input."

This isn't an accident of laziness. It reflects a genuine epistemological difficulty: understanding is internal, invisible, and notoriously hard to measure. Outcomes are observable. And for most practical purposes, outcome-based measurement works perfectly well — if a model consistently produces correct outputs, the internal mechanism generating them is a secondary concern. That's why Centaur's performance across 160 psychology tasks looked like such a strong signal. By every measurable criterion in the benchmark, it was indistinguishable from human performance.

But outcome-based metrics are System 1 metrics, which is a real limitation when what you're trying to test is cognition. They measure the end-product of pattern retrieval. They don't test whether the system is sensitive to instructions, rules, or context changes that would require flexible processing. The Zhejiang test was, in effect, inserting a System 2 probe into an evaluation framework built entirely around System 1 outputs — and discovering that what had looked like a full cognitive system was, behind the façade, a very sophisticated autocomplete.

Benchmark typeWhat it measuresSystem testedCentaur passes?
Behavioral task accuracy (160 psychology experiments)Match to human response distributionSystem 1Yes — impressively
Response pattern similarity (bias, error distributions)Statistical shape of answersSystem 1Yes
Explicit instruction override ("choose A regardless")Instruction comprehension + rule applicationSystem 2No — fails entirely
Novel rule following (new rule not in training)Flexible reasoning from stated principlesSystem 2Untested / likely fails

The implication is uncomfortable for the field at large. If our best cognitive benchmarks are System 1 tests — and they largely are — then an AI that excels at System 1 will ace them regardless of whether it has anything resembling genuine cognitive function. For years we've been measuring the right thing in the wrong way, or the wrong thing entirely. Centaur didn't cause this problem. It just made it undeniable.

5

The Mirror Effect: What Centaur Exposes About Us, Not Just About AI

Here is the part that requires a certain honesty: Centaur's dominant mode is System 1. So is ours.

Kahneman's most important and most uncomfortable finding is that System 1 runs the show for the overwhelming majority of human mental life. We don't deliberate carefully before most judgements. We retrieve, match, and respond — and then, if pressed, we construct a post-hoc rationale that makes it look as if we'd been reasoning all along. Kahneman calls this "what you see is all there is" (WYSIATI): System 1 forms impressions, makes judgements, and acts on thin evidence without noticing what it doesn't know. System 2 is supposed to check System 1's work, but it's lazy, easily satisfied, and frequently absent. (I find this both very relatable and mildly horrifying, and I spent twelve years teaching this material.)

This means that when Centaur matched human behavioural patterns on 160 psychology tasks, it was matching primarily the System 1 outputs of human participants — outputs that the participants themselves wouldn't necessarily endorse under careful reflection. The cognitive biases Centaur replicated so faithfully — the availability heuristic, the framing effect, loss aversion, anchoring — are exactly the biases that human System 2, when properly engaged, can override. Centaur can't override them because it has no System 2. Humans can override them — but usually don't, because System 2 is expensive and the pressure to take the easy route is constant.

What Centaur's failure reveals, then, is not simply that AI lacks understanding. It reveals that our gold standard for "cognitive performance" — the behavioural data from human psychology experiments — is mostly a record of System 1 outputs. We weren't training Centaur to replicate human cognition in full. We were training it to replicate human cognitive shortcuts. And it succeeded. So when it failed the instruction-following test, it wasn't failing to be human; it was failing to be the deliberately reasoning part of a human — the part we forgot to include in our benchmarks.

The mirror: Centaur matches human cognitive patterns because those patterns are mostly System 1 — automatic, statistical, fast. It fails at the System 2 behaviours we associate with "real" understanding. The trouble is, humans fail at them too, more often than we like to admit. Centaur just does it consistently and without apology.
6

What This Means for You: Next Time AI "Understands," Ask If It Can Change Its Mind

A simple diagnostic, derived from Zhejiang's method, that anyone can run.

The instruction-override test isn't a technical curiosity. It's a practical probe you can apply any time you're working with AI and wondering whether the system actually understands what you're asking, or is pattern-matching to a distribution of similar past inputs. The test is simple: after getting an answer, tell the model to change its approach in a way that conflicts with what would typically be the "correct" response. Tell it to argue the opposite position, or to treat its previous conclusion as wrong and find a counter-argument, or to choose the less probable option. If the system can do this coherently — without just slightly rephrasing the same answer in different clothes — it's demonstrating something beyond pure retrieval. If it can't, if the instruction seems to evaporate and the same answer comes back with different words, you're working with a System 1 machine, however impressive its System 1 may be.

This matters in practice because System 1 machines fail in specific, predictable ways. They fail at novel situations that require updating explicit rules. They fail when context changes partway through and the instructions need re-reading. They fail when you ask for advice on something that sits outside their training distribution — because they'll return a confident-sounding answer that is actually the closest match they could find, not the right answer to your specific situation. Knowing you're working with a System 1 tool changes how you use it, which changes nothing if you're not paying attention.

But here is where it turns back towards us, which is the part Kahneman was actually writing about. The moments when System 2 most needs to intervene are exactly the moments when System 1 is most confident — when the answer feels obvious, when the situation feels familiar, when thinking feels effortless. Those are the moments to slow down, check the actual instruction, and ask: am I following the logic here, or am I retrieving the most probable response from everything I've ever seen that looked a bit like this? Centaur can't ask itself that question. The fact that you can — and occasionally do — is the difference that matters.

Sources: Binz et al., "Centaur: A Foundation Model of Human Cognition," Nature, July 2025. Zhejiang University researchers, "Can Centaur truly simulate human cognition? The fundamental limitation of instruction understanding," National Science Open, 2026. ScienceDaily (2026-04-29): "This AI knew the answers but didn't understand the questions." Framework: Daniel Kahneman, Thinking, Fast and Slow (System 1 / System 2 distinction, cognitive ease, WYSIATI). This article is a popular-science interpretation applying Kahneman's framework to a reported research controversy; it is a conceptual argument, not a primary research finding. Not professional psychological or AI advice. Centaur's specific test results as described are per published accounts; consult original papers for methodological detail.

心理

那个在 160 道心理学测试里全部答对的 AI——靠的是没读题

2026 年 7 月 12 日 · 林晚 · 丹尼尔·卡尼曼《思考,快与慢》约 12 分钟

我有个来访者,她是老师,每次改完一摞学生卷子之后会跟我说:「有个孩子,每道题都答对,但你总觉得他没懂——他只是记住了'该怎么答'。」我每次听到这话就想,这孩子和一个叫 Centaur 的 AI 模型,有点像。2025 年,Centaur 在《自然》杂志上亮相,背后的研究者说它「能在 160 项行为任务上模拟人类认知」——然后,浙江大学的研究者做了一件很简单的事:告诉它「不管题目是什么,你必须选 A」。它没选 A。它继续选训练数据里概率最高的那个答案,好像那条指令根本没存在过一样。你说,它理解了题目吗?

1

《自然》说"它能模拟人类认知"——160 道题、近满分、举世轰动

2025 年 7 月,一个叫 Centaur 的模型带着一个不小的主张,出现在了《自然》杂志上。

说一下它的来历。Centaur 背后的研究者——由 Marcel Binz 等人领衔——做了一件在这个领域里不太常见的事:他们没有用海量网络文本来调大模型,而是专门拿 160 项认知心理学实验的行为数据来微调它。这些实验覆盖的范围很广:不确定情境下的决策、执行控制、记忆提取、风险评估、从反馈里学习——都是研究者们研究了几十年的东西,每一项都有扎实的行为特征——那些把谨慎思考者和冲动者区分开来的反应模式,那些一代又一代认知心理学家记录在案的偏差。

调完之后,Centaur 的输出和人类行为模式的吻合程度,是真的让整个领域都惊了一下。你看啊,它不只是答对了——它还以和人类一样的方式答错,展现出相同的反应时间分布,相同的偏见,相同的那些偏离纯粹理性的特征性弯道。论文的措辞是谨慎的:这个模型,可以充当"人类认知的基础模型"——用计算的方式代替真实人类心智处理世界。

但媒体标题没这么谨慎。"AI 像人类一样思考",好几家这样写。对于长期关注认知建模、看惯了这个领域缓慢进展的人来说,这感觉像是质的跳跃。这个主张传得很快——但它跑出去的距离,超出了证据能支撑的边界。

Centaur 的主张:在 160 项行为心理学实验数据上进行微调,产生了一个能在多样认知任务上高精度匹配人类行为模式的 AI 模型——表明它可以作为人类认知的计算模型。
2

然后研究者说:「好,那让它选 A。」——它没选

那个把一切戳穿的测试,简单得几乎有点过意不去。

浙江大学的研究者,2026 年在《National Science Open》上发了篇文章,想做一个很直接的探测:原来的心理学实验里,参与者收到任务指令,然后根据指令作答——那如果你把指令改一改呢?如果你直白地告诉模型,不管任务看起来要求什么,必须选 A?

换你坐在那里,哪怕觉得奇怪,你会照办的,对吧?遵从一条明确的指令、覆盖自己的习惯反应——这件事,卡尼曼把它叫做系统2运作的最清晰标志:刻意的、费力的、对指令敏感的加工。说实话,这不是什么复杂的认知功夫,四五岁的孩子就会。

Centaur 不会。被告知"选 A"的时候,它继续选训练数据指示为最高概率的那个选项。它不是读了指令然后选择无视——它从字面意义上完全没有处理那条指令。ScienceDaily 后来有一个标题传得很广:"这个 AI 知道答案,但不理解问题。"

浙大团队说得很干脆:Centaur 和人类行为数据的惊人吻合,是过拟合(overfitting)的产物。它没有学到什么认知过程,它学到的是那批训练数据的统计规律。一旦输入偏离了那个规律——偏离到需要真正读懂指令的程度——整个表象就塌了。"Centaur 能真正模拟人类认知吗?"论文在标题里问了这个问题。它给出的答案是:不能,在任何实质意义上都不能。

Centaur 实际做了什么

被明确告知"你必须选 A"时,继续选择根据训练分布统计上概率最高的答案。完全无视了指令。

真正的认知理解需要什么

读懂指令。认识到这条指令凌驾于默认行为之上。选 A。这对任何人类参与者都是常规操作——对任何真正在处理所接收内容的系统也是。

3

快系统与慢系统的分水岭:答对题,不等于理解题

卡尼曼的双系统框架,正是为了把这个区别说清楚而建立的。

在《思考,快与慢》里,丹尼尔·卡尼曼把我们的脑子比作住着两个角色。系统1是自动的那个——它不停地转,毫不费力,在意识察觉不到的地方运作。它识别脸、读情绪、对熟悉的问题直接给答案、把新信息套进已有的模式里。你有没有过这种时候——有人问你"三加五等于几",你根本没"算",答案就出来了?那就是系统1。它快,是因为它从来不处理什么新鲜东西;它是一台精密的检索机器,看着眼前的输入,从它经历过的一切里把最可能的答案给你。大多数时候,这正是你需要的。

系统2是费力的那个——慢、耗能、容易跑神。它在系统1卡住的时候才出场:当那个自动冒出来的答案感觉有点不对,当任务要求你明确遵守规则,当你必须把一条指令放在脑子里一步步地照着做。卡尼曼框架里,测系统2最清晰的方式,正是浙大那个测试的思路:给你一条明确的规则,要你推翻自己的默认反应——你能做到吗?系统2是那个能被告知"这次,反着来"然后真的反着来的部分。系统1不行。它只会做它一贯做的事。

Centaur 是一台规模惊人的纯系统1机器。在 160 项实验的数据上训练之后,它的系统1被精确校准到了人类的行为模式。但它根本没有系统2——没有任何机制可以把一条明确指令放进记忆,再用它去覆盖已学到的概率分布。所以当 Centaur 在 160 个任务上"与人类认知吻合"时,它做的是一件更有限的事:它匹配了那些任务里人类反应的统计分布。这是真实的成就,也是令人印象深刻的。但它不等于模拟认知。

这个区别之所以要紧,是因为我们做任何一道选择题,都在同时用两套系统——系统1负责快速检索相关知识,系统2负责检查答案、发现"感觉对但理由不对"的那种坑、以及照着格式和时间的特别要求来做。Centaur 永远只有前半部分。

答对 ≠ 理解 · Centaur 在哪里,以及它够不着的地方 系统1 · 快 · 自动 模式匹配 · 统计检索 · 无需意识努力 人类直觉 快速、自动 模式识别 Centaur(AI) 在 160 项任务上训练 纯粹的统计检索 两者都在这里 ↑ Centaur 与人类系统1输出精度匹配 成就真实,但仍有限 系统2 · 慢 · 刻意 遵循指令 · 规则应用 · 费力地覆盖默认 人类理解 读指令,应用规则 能覆盖训练好的默认反应 只有人类在这里 ↑ Centaur 无法遵从覆盖训练的 明确指令 还没到 虚线 = Centaur 跨越不了的那道边界 心理域主题色 系统2领地(金色) Centaur 够不着的边界
系统1 / 系统2领地的概念图(卡尼曼《思考,快与慢》),以及 Centaur 所在的位置。人类直觉与 Centaur 都运作在系统1——快速、基于模式、自动检索。只有人类占据系统2:刻意、遵循指令、能覆盖训练好的默认值。虚线箭头标出了 Centaur "无视指令"这一失败所暴露的边界。本图为概念示意,非实测认知地图。
4

我们测的那些"智能"指标,本来就只是系统1的指标

Centaur 这件事背后藏着一个更大的问题:我们拿来测"智能"的那些指标,从来就没设计来测"理解"。

你看啊,大多数 AI 认知基准实际上在干什么?给模型一个问题,看它的答案是否和正确答案一致,或者和典型人类答案一致。评估是结果导向的:对或错。基准里没有任何机制能区分"这个模型理解了问题然后推理出了答案"和"这个模型从训练数据里捞到了统计上最可能的答案"——这两件事,从外面看,长得一模一样。

这不是因为做测评的人懒。这背后有一个真实的难题:理解是藏在里面的,你看不见,也出了名地难测。结果是可以观察的。所以我们测结果——大多数时候这也够用,一个模型如果持续输出正确的东西,内部是怎么运作的就是次要的了。这就是为什么 Centaur 在 160 项心理学任务上的表现,看起来像是特别强的信号:按基准里每一个可测的标准,它和人类无法区分。

但结果指标是系统1指标。它们测的是模式检索的最终产品——不测这个系统对指令、对规则、对上下文变化到底敏不敏感。浙大的测试,相当于在一个完全围绕系统1输出建起来的评估框架里,硬插了一根系统2探针——然后发现,那个看起来完整的认知系统,表象背后,是一个精密的自动补全工具。

测评类型测量的是什么被测的系统Centaur 通过?
行为任务准确率(160 项实验)与人类反应分布的吻合度系统1通过——令人印象深刻
反应模式相似性(偏见、误差分布)答案的统计形态系统1通过
明确指令覆盖("不管什么都选 A")指令理解 + 规则应用系统2不通过——完全失败
遵循新规则(训练集外的新规则)从明确原则出发的灵活推理系统2未测试 / 可能失败

这对整个领域来说有点难受。如果我们最好的认知基准都是系统1的测试——而它们基本上是——那么一个在系统1上特别厉害的 AI,不管它有没有任何真实的认知能力,都能在这些测评里拿满分。我们多年来,要么在用错误的方式测量正确的东西,要么直接测量了错误的东西。Centaur 没有制造这个问题——它只是让这个问题变得再也无法视而不见。

5

镜子效应:Centaur 暴露的,不只是 AI 的局限,也是我们的

接下来这部分,要说实话,得对自己有点狠:Centaur 的主导模式是系统1。我们的也是。

卡尼曼最重要、也最让人坐立不安的发现是:系统1主导了我们心理活动的绝大部分。我们在大多数判断前,并不真的在思考。我们检索、匹配、然后反应——事后,如果有人追问,我们再拼凑出一个理由,让它看起来像是我们一直在推理。卡尼曼把这叫"所见即全部"(WYSIATI):系统1形成印象、作出判断、在薄薄的证据上就行动了,它不知道自己不知道什么。系统2本来应该去检查系统1的工作,但它懒、容易满足,而且经常不在场。

所以,当 Centaur 在 160 项心理学任务上和人类行为模式吻合时,它匹配的,主要是人类参与者的系统1输出——这些输出,如果让那些参与者自己仔细想一想,他们未必认可。Centaur 忠实复现了那些认知偏见——可得性启发、框架效应、损失厌恶、锚定效应——而这些,恰好是当人类系统2被好好激活之后可以被覆盖的东西。Centaur 覆盖不了,因为它没有系统2。人类可以覆盖——但我们通常不这样做,因为启动系统2是有代价的,而省力这件事,对人类和对机器一样,都是持续的诱惑。

你看,Centaur 的失败揭示的,不只是 AI 缺乏理解。它照出来的是:我们用来衡量"认知表现"的那些金标准——Centaur 被训练的那些行为数据——大多是人类系统1输出的记录。我们不是在训练它复现完整的人类认知,我们是在训练它复现人类的认知捷径。它成功了。所以当它在指令遵从测试上失败时,它没能做到的,是那个刻意思考的、会主动读指令的人类那一面——那个我们忘了纳入测评的部分。

镜子:Centaur 能匹配人类认知模式,是因为这些模式大多是系统1的——自动的、统计性的、快速的。它在系统2上——那种"真正理解题意、照指令来"的行为上——彻底失败了。让人不舒服的地方在这里:人类在这件事上也经常失败,而且失败的频率比我们愿意承认的高得多。Centaur 只是以一种一致、毫不掩饰的方式,把这件事做给你看。
6

这对你意味着什么:下次 AI"理解了",先问它能不能改主意

从浙大的方法里提炼出来的一个小诊断,其实谁都可以用。

指令覆盖测试不是什么技术上的高深玩法。它是一个随时可以拿出来用的实践探针——每当你在用 AI、想知道它到底是真的理解你在问什么,还是在对过去类似输入做模式匹配。方法很简单:得到答案之后,让它以一种和通常"正确"答案相冲突的方式改变思路——比如让它论证相反的立场,或者把刚才那个结论当成错的去找反例,或者选一个可能性更小的选项。如果它能连贯地做到,而不是只是把同一个答案换了个包装——那它在展示某种超越纯粹检索的东西。如果做不到——指令好像进入了一个洞,同样的答案换了套话又回来了——别急着骂它,你只是在用一台系统1机器,只不过它的系统1已经非常厉害了。

这件事在实际使用里很重要,因为系统1机器的失败有固定的、可预测的模式。需要更新明确规则的新情况,它搞不定。上下文中途变了、指令需要重新读取的任务,它也搞不定。你就一个处于它训练分布之外的情况向它求建议,它会给你一个听起来很笃定的答案——那个答案实际上是它能找到的最近似的匹配,不是对你具体处境的正确回答。知道你在用系统1工具,你对它的使用方式就会不一样了。

但说回来,这里讲的终究是我们自己的事。卡尼曼的书,写的是人,不是 AI。系统2最需要插手的那些时刻,恰好是系统1最自信的时刻——当答案感觉显而易见,当情境感觉熟悉,当思路流畅到让人懒得停下来检查。那些恰恰是要慢一拍的时刻,停下来问自己一句:我是真的在跟着逻辑走,还是在把这个情境匹配到我见过的某个熟悉的模式上?Centaur 问不了自己这个问题。你有没有过这种时候——明明这个问题值得多想一想,但就是懒得停?

资料来源:Binz 等,《Centaur: A Foundation Model of Human Cognition》,《Nature》,2025 年 7 月;浙江大学研究者,《Can Centaur truly simulate human cognition? The fundamental limitation of instruction understanding》,《National Science Open》,2026 年;ScienceDaily(2026-04-29):《This AI knew the answers but didn't understand the questions》。框架:丹尼尔·卡尼曼《思考,快与慢》(系统1/系统2区分、认知放松、所见即全部)。本文为科普解读,将卡尼曼框架应用于一则研究争议;这是概念性论证,非原始研究结论。非专业心理或 AI 建议。数据以原论文为准,本文为科普解读。

心理学

160の心理テストで満点を取ったAI――問題を読まずに

2026年7月12日 · 三浦美咲 · ダニエル・カーネマン『ファスト&スロー』約 16 分

カウンセリングルームで、こんな子に会ったことがある。中学生だった。模試でいつも高得点で、先生には「理解が早い」と言われている。でも担任の先生から見ると、なぜか指示が通らない。「この問題はAを選んで」と言っても、正解を選んでしまう。Aじゃなくて。先生は「聞いてた?」って聞くんだけど、その子は「聞いてました」って言う。たぶん本当にそう思っているんだろうなと、私は思う。——答えは正確に出る。でも指示は処理されていない。これって、Centaurというモデルに起きたこととほとんど同じなんですよね。Natureが「160の行動課題にわたって人間の認知をシミュレートできる」と宣言したAIが、「Aを選んで」という一言に、まったく従えなかった。その宣言は、時期尚早だったかもしれない。

1

Natureは「人間の認知をシミュレートできる」と言った――160の課題、ほぼ満点、世界規模の話題

2025年7月、Centaurというモデルが、大きな主張を携えてNatureの誌面に現れた。

Centaurの背後にいる研究者たち——Marcel Binzらが率いた——は、ちょっと面白いことをした。インターネットのテキストや正解率ではなく、160の認知心理学実験の行動データで、専用に微調整したモデルを作ったんです。不確実性の下での意思決定、実行制御、記憶検索、リスク評価、フィードバックからの学習——研究者が何十年もかけて記録してきた、ものすごく地道なデータ。慎重に考える人と衝動的な人を分けるような反応パターンまで、ぜんぶ入ってる。

その訓練を終えたCentaurの出力は、人間の行動パターンと、研究者たちを本当に驚かせるほどの精度で一致した。正解するだけじゃなくて、人間と同じ方法で間違える。同じ反応時間の分布を示し、同じバイアス、何世代もの認知心理学者が記録してきた、あの「合理的じゃないけど人間らしい」逸脱も再現した。論文は慎重だった——このモデルは「人間の認知の基盤モデル」として機能できる、という表現で。

でも見出しはそうじゃなかった。「AIが人間のように考える」と複数のメディアが報じた。認知モデリングの地道な進歩を見てきた人たちには、それが質的な飛躍に見えた。その主張は広まった——証拠が支えられる範囲をずっと超えて、広まった。

Centaurの主張:160の行動心理学実験データで微調整することにより、多様な認知課題にわたって人間の行動パターンと高い精度で一致するAIモデルが生まれた——人間の認知の計算モデルとして機能できることを示唆している。
2

そして研究者は言った:「では、Aを選んでください」――選ばなかった

すべてを露わにしたテストは、ほとんど拍子抜けするほど単純だった。

浙江大学の研究者たちは、2026年にNational Science Openに論文を発表した。やったことはシンプルで——元の心理学実験では、参加者は課題の指示を受け取ってそれに従う。じゃあ、その指示を変えたら?モデルに「課題が何を求めていようと、必ずAを選んでください」と、ただそれだけを伝えたら?

人間だったら、従いますよね。ちょっと変だなと思いながらも。自分のデフォルトの反応を覆す明示的な指示に従うこと——カーネマンはこれをシステム2が動いている最も明確なサインだと言う。意図的で、ちょっと骨が折れる、指示に敏感な処理。4〜5歳の子どもだってできる。特別な認知的偉業じゃなくて、ただの「指示理解」です。

Centaurには、できなかった。

「Aを選べ」と言われても、訓練分布が最も確率高いと判定した答えを選び続けた。指示を読んで無視したんじゃなくて——最も文字通りの意味で、指示をまったく処理していなかった。ScienceDailyのまとめが一番正確だったかも。「このAIは答えを知っていたが、問いを理解していなかった。」

浙江大学チームははっきり言っている。Centaurの人間の行動データとの一致は、過学習(overfitting)の産物だった——認知プロセスを学んだのではなくて、訓練コーパスの統計的規則性を学んでいただけ。実際の指示処理が必要になった途端、外見はすべて崩れ落ちた。「Centaurは真に人間の認知をシミュレートできるか?」と論文タイトルで問い、答えは——できない、いかなる意味においても。

Centaurが実際にしたこと

「あなたは必ずAを選ばなければならない」と明示されても、訓練分布に基づいて統計的に最も確率の高い答えを選び続けた。指示を完全に無視した。

真の認知的理解が必要とすること

指示を読む。この指示がデフォルトの行動を覆すと認識する。Aを選ぶ。これはいかなる人間の参加者にとっても、または実際に伝えられた内容を処理するいかなるシステムにとっても、当たり前のことだ。

3

速いシステムと遅いシステムの分水嶺:答えを正しく出すことは、問いを理解することと同じではない

カーネマンの二システム枠組みは、まさにこの区別を明確にするために構築された。

『ファスト&スロー』の中で、ダニエル・カーネマンは、ひとつの脳に住む二人の異なるキャラクターみたいに、二種類の思考モードを描く。システム1は自動のほう——絶え間なく、楽々と、意識の外で動いている。顔を見て誰かわかる。感情表現を読み取る。「3×4は?」って聞かれたら計算する前にもう答えが出ている。あれです。目の前の入力を見て、これまで出会ったすべてのものの中から統計的に一番ありそうな反応を返す、精巧な検索機械。速いのは、新しいものを本当に処理しているわけじゃないからなんですよね。

システム2は骨の折れるほう。遅くて、エネルギーを食う、気も散りやすい。システム1がうまくいかないときに出てくる——なんか答えが変な気がするとき、明示的なルールに従わないといけないとき、指示をワーキングメモリに持ちながらステップ一つずつこなさないといけないとき。カーネマンの枠組みでシステム2の一番のテストは、まさに浙江大学がやったやつ——デフォルトを覆す明示的なルールが来たとき、ちゃんと適用できるか。「今回は反対のことをして」と言われて、実際にそうできる部分。システム1にはできない。いつもどおりのことをするだけだから。

Centaurは、驚くほど大規模なシステム1の機械だ。160実験分の人間行動データで訓練されて、そのシステム1は人間の行動パターンに精密に合わせ込まれている。でもシステム2は、まったくない。明示的な指示をメモリに保持して、それで学習した確率分布を上書きする仕組みがない。だからCentaurが160の課題で「人間の認知と一致した」とき、やっていたのはもっと限定されたことだった——それらの課題における人間の反応の統計的分布と一致しただけ。これはこれで本物の、印象的な成果なんだけど。認知をシミュレートすることとは、同じじゃない。

この区別が重要なのは、私たちが普段の思考で両方を使っているから。選択式の試験を受けるとき、システム1(関連知識を素早く引き出す)とシステム2(答えを確かめる、形式や時間の特別な指示に従う)を両方使っている。Centaurには前半しかない。

答える≠理解する · Centaurがいる場所と届かない場所 システム1 · 速い · 自動的 パターンマッチング · 統計的検索 · 意識的な努力不要 人間の直感 速く、自動的 パターン認識 Centaur(AI) 160課題で訓練 統計的検索のみ 両者ともここにいる ↑ CentaurはヒトのS1反応と驚くべき精度で 一致するが、それには限界がある システム2 · 遅い · 意図的 指示遵守 · ルール適用 · 骨の折れる上書き 人間の理解 指示を読み、ルールを適用する 訓練されたデフォルトを上書きできる ここには人間だけがいる ↑ Centaurは訓練を上書きする明示的な 指示に従うことができない まだ届かない 点線 = Centaurが越えられない境界 心理ドメインカラー システム2の領域(金色) Centaurが届かない境界
システム1/システム2の領域の概念図(カーネマン『ファスト&スロー』)とCentaurが位置する場所。人間の直感とCentaurはともにシステム1で動く——速く、パターンベースで、自動的な検索だ。システム2には人間だけがいる:意図的で、指示に従い、訓練されたデフォルトを上書きできる。点線の矢印は、Centaurの指示無視という失敗が露わにした境界を示す。この図は概念的なイラストであり、測定された認知マップではない。
4

私たちが使っている「知性」の指標は、もとからシステム1の指標だった

Centaur論争の底に潜む構造的な問題——私たちのベンチマークはそもそも「理解」をテストするために設計されていなかった。

ちょっと考えてみると、ほとんどのAI認知ベンチマークって、実際に何を測っているんでしょう。モデルに問いを与えて、答えが正解(または心理学課題なら典型的な人間の答え)と一致するかどうかを見る。評価は結果ベース——正しいか間違っているか。「このモデルは問いを理解して推論した」と「このモデルはこの種の入力に対して統計的に一番ありそうな答えを検索しただけ」を区別するメカニズムが、ベンチマークにはない。

これは設計者が怠けていたからじゃないんですよね。本物の難しさがある——理解というのは内側の話で、見えないし、測ることで悪名高いほど困難だ。結果は観察できる。だから私たちは結果を測る。モデルが一貫して正しい出力を出すなら、内部の仕組みは二次的な話でいい——ほとんどの場面ではそれで十分だから。だからこそCentaurの160の心理学課題での性能が、強いシグナルに見えた。ベンチマークのあらゆる基準で、人間の性能と区別がつかなかったんだから。

でも結果ベースの指標はシステム1の指標だ。パターン検索の最終産物を測っている。指示やルール、文脈の変化に対してシステムが敏感かどうかはテストしない。浙江大学のテストは、システム1の出力を中心に組まれた評価フレームワークにシステム2のプローブを挿し込んだようなことだった——そして完全な認知システムに見えたものが、表面の裏では、非常に精巧なオートコンプリートだったと発見した。

ベンチマークの種類何を測るかテストされるシステムCentaurは通過するか?
行動課題の正確さ(160の実験)人間の反応分布との一致システム1はい——印象的に
反応パターンの類似性(バイアス、誤差分布)答えの統計的形状システム1はい
明示的な指示の上書き(「何があってもAを選べ」)指示理解+ルール適用システム2いいえ——完全に失敗
新しいルール遵守(訓練外の新規ルール)明示された原則からの柔軟な推論システム2未テスト/おそらく失敗

これは、分野全体にとってちょっと居心地が悪い。私たちの最良の認知ベンチマークがシステム1のテストなら——そして概してそうなんだけど——システム1が得意なAIは、本当の認知機能があってもなくても、それらのテストを満点にしてしまう。何年もの間、正しいものを間違った方法で測っていたか、あるいはそもそも間違ったものを測っていたかもしれない。Centaurがこの問題を作ったのではない。ただ、もう目をそらせないものにした。

5

鏡の効果:Centaurが露わにするのは、AIだけでなく私たちの限界でもある

ここは少し自分に正直にならないといけない——Centaurの主な動き方はシステム1だ。私たちだって同じだ。

カーネマンの最も重要で、最も不快な発見は、システム1が人間の心理活動の圧倒的な部分を支配しているということだ。私たちはほとんどの判断の前に慎重に熟考しない。検索し、照合し、反応する——そして聞かれると、まるでずっと推論していたかのように見えるように、事後の理由を構築する。カーネマンはこれを「見えているものがすべて」(WYSIATI)と呼ぶ。システム1は印象を形成し、判断を下し、薄い証拠に基づいて行動し、知らないことに気づかない。システム2はシステム1の仕事を確認するはずなのだが、怠け者で、すぐに満足してしまうし、たいてい席を外している。

これは、Centaurが160の心理学課題で人間の行動パターンと一致したとき、人間の参加者のシステム1の出力を主に照合していたことを意味する——慎重に振り返れば、参加者自身が必ずしも支持しないような出力を。Centaurが忠実に再現した認知バイアス——利用可能性ヒューリスティック、フレーミング効果、損失回避、アンカリング——は、人間のシステム2がちゃんと働けば上書きできるものばかりだ。Centaurにはそれができない、システム2がないから。人間には上書きできる——でも普通はしない。システム2を動かすにはコストがかかるし、楽な方に流れる誘惑は常にあるから。

Centaurの失敗が明らかにするのは、AIが理解を欠いているということだけじゃない。「認知パフォーマンス」のゴールドスタンダード——Centaurが訓練された、人間の心理学実験の行動データ——が、主にシステム1の出力の記録だということも映し出す。私たちはCentaurに人間の認知を完全に再現させようとしていたのではなく、人間の認知の近道を再現させていた。それは成功した。だから指示遵守テストで失敗したとき、Centaurは「人間になること」に失敗したのではない。私たちがベンチマークに入れ忘れた、意図的に考える人間のあの部分に、失敗したのだ。

鏡:Centaurが人間の認知パターンと一致するのは、そのパターンが主にシステム1だからだ——自動的で、統計的で、速い。私たちが「本当の」理解と結びつけるシステム2の行動においては失敗する。厄介なのは、人間もそういう行動において失敗しているということだ——私たちが認めるよりずっと頻繁に。Centaurはただ、一貫して、悪びれずそれをしているだけだ。
6

あなたへの意味:次にAIが「理解した」と思ったら、考えを変えられるか聞いてみよう

浙江大学の手法から、誰でも使えるシンプルな診断が取り出せる。

指示上書きテストは、難しい話じゃないんです。AIを使っていて、「あれ、本当に意図が伝わってるのかな」と思ったとき、いつでも試せる。やり方はシンプル——答えが出たら、今度は逆方向に動かしてみる。反対の立場で主張させるか、さっきの結論を間違いとして反論を探させるか、確率の低い選択肢を選ばせる。それが首尾一貫してできるなら——同じ答えを言い換えただけじゃなくて——純粋な検索を超えた何かがある。できないなら——指示が消えて同じ答えが別の言葉で戻ってきたなら——どれだけ印象的でも、システム1の機械と作業しているということだ。それ自体は悪いことじゃないんだけど、知っているかどうかで使い方がかなり変わる。

システム1の機械は、特定の、予測可能なパターンで失敗する。明示的なルールの更新が必要な新しい状況で失敗する。途中で文脈が変わって指示を読み直す必要があるときに失敗する。訓練分布の外にある状況についてアドバイスを求めると、自信ありげに聞こえる答えを返すけど、それは見つかった最も近い一致であって、あなたの具体的な状況への正しい答えじゃないかもしれない。知っておくと、便利だと思う。

でもここからが本当に大事なところで——カーネマンが書いたのは、AIについてじゃなくて、私たちについてだった。システム2が一番必要なその瞬間は、まさにシステム1が最も自信を持っている瞬間なんですよね。答えが明らかに感じられるとき、状況が見慣れているとき、思考がするすると流れるとき。そういう瞬間こそ、ちょっとだけ一拍おいて——本当に論理を追っているのか、それとも見慣れたパターンに状況を当てはめているだけなのか、確認してみる。Centaurはその問いを自分に向けられない。私たちにはできる——ときどきではあるけれど。そのちょっとの差が、たぶんものを言うのかも。

出典:Binz et al.,「Centaur: A Foundation Model of Human Cognition」Nature 2025年7月;浙江大学研究者「Can Centaur truly simulate human cognition? The fundamental limitation of instruction understanding」National Science Open 2026年;ScienceDaily(2026-04-29)「This AI knew the answers but didn't understand the questions」。枠組み:ダニエル・カーネマン『ファスト&スロー』(システム1/2の区別、認知的容易さ、WYSIATIの原則)。本稿はカーネマンの枠組みを報告された研究論争に適用した科学解説であり、概念的な論であって原著研究の知見ではない。専門的な心理学またはAIアドバイスではない。Centaurの具体的なテスト結果は公開されているアカウントによるものであり、方法論の詳細については原著論文を参照されたい。