A Model That Only Knows Your Data Can Never Leave It
June 21, 2026 · Wu Jun, The Beauty of Mathematics~6 min read
Roughly nine out of ten of the world's strongest magnets pass through one country before they reach your earbuds, your EV motor, your wind turbine. China runs about 85–90% of rare-earth separation and roughly 90% of the magnet manufacturing that follows. That's not a chemistry fact — it's a chokehold. So when a lab says it has an AI plan to design a powerful magnet with no rare earths in it at all, the obvious question is: real escape route, or wishful press release? The honest answer turns on a quiet distinction most coverage skips — the difference between a model that has seen a lot and a model that understands something.
The chokehold nobody voted for
Today's benchmark magnet, NdFeB, owes its punch to neodymium and a few other rare-earth elements. Mine them anywhere you like; the separation step — pulling one chemically near-identical metal from its neighbors — is where the bottleneck bites, and that step lives overwhelmingly in China. A trade spat, an export license, a quiet quota change, and the supply of the thing inside every motor and hard drive tightens at someone else's discretion. The cleanest fix isn't a friendlier supplier. It's a magnet that doesn't need the rare earth in the first place. The trouble is that nobody has found one strong enough, and the search space — every possible alloy, every crystal arrangement — is far too vast to test by hand or by luck.
Two ways to be smart about a haystack
In 2025, Prashant Singh, a researcher at Ames National Laboratory, published a roadmap (in Advanced Functional Materials) for letting AI do the hunting. And he drew a line worth quoting, because it's the whole game: "If you just use the data to train your models, you will get only predictions within the range of information you have. But once you understand the physics of what controls specific properties, then you and your agentic tools or AI frameworks can search arbitrary material space." Read that twice. A model fed only examples learns the shape of what it was shown and stays politely inside it — ask for something beyond the fence and it guesses by stretching the nearest pattern. A model that carries the physics of why a material behaves as it does isn't fenced by its examples. It can reason out toward materials no one has measured, because it holds the rule, not just the cases.
A model trained only on examples can predict inside the range of data it has; once the model carries the physics that controls a property, it — and the AI tools built on it — can search arbitrary material space. The catch: this is a roadmap, not a discovered magnet (no synthesis, scaling or cost yet), while China controls ~85–90% of rare-earth separation. Framework: Wu Jun, The Beauty of Mathematics. Per Prashant Singh / Ames National Laboratory, Advanced Functional Materials 2025; skeptics per rareearthexchanges.com. Popular-science reflection.
This is the oldest lesson in computing
Singh just rediscovered, in metallurgy, what Wu Jun spends a whole book on in The Beauty of Mathematics: the right model of a problem beats more raw data about it. Early machine translation tried to brute-force language with dictionaries and hand-written grammar rules and went nowhere; it leapt forward when engineers modeled language as probability — a structure that captured how words actually hang together. Across his examples the moral repeats: when you find the mathematics that mirrors the real structure of a problem, you stop memorizing answers and start being able to derive them. Call the failure mode the data ceiling — the invisible wall where a pattern-matcher, no matter how much you feed it, can only ever hand back rearrangements of what it already saw. Physics is how you climb over the wall instead of repainting it. A magnet's strength isn't a vibe to be interpolated; it's set by how electron spins line up in a crystal lattice, and that is describable, in equations. Hand a model those equations and "search every alloy" stops being a fantasy and becomes a calculation.
Now the part the headlines skip
Here's where a builder has to stay honest, because the gap between "promising framework" and "problem solved" is exactly where hype lives. Read the fine print and skeptics — among them the trade outlet rareearthexchanges.com — make the load-bearing point: this paper does not report the discovery of a rare-earth-free permanent magnet. It's a roadmap. No new magnet has been synthesized, scaled, or shown to compete on cost. On the famous hype curve this sits near the very start, the "Technology Trigger" — the moment of maximum excitement and minimum proof. And there's a subtler catch that Wu Jun's own framework predicts: a physics-grounded model is only as trustworthy as the physics inside it, and that physics is itself anchored to experimentally measured numbers. Feed it shaky measurements and your elegant equations will extrapolate confidently into nonsense. (Worth flagging, too: a separate AI tool called DuctGPT, aimed at alloy ductility for fusion and aerospace, has been floating around the same conversation — it did not discover a magnet either.) The model can reach arbitrary material space; whether what it finds out there is real still has to be checked at a bench.
What this means for you
Strip it to the part you can actually use, in a lab or anywhere else you build with AI. More data makes a system better at the world it already knows; a better model is what lets it act in a world it hasn't seen. When someone sells you an AI that "learned from millions of examples," the sharp question isn't how many examples — it's whether the thing grasps the underlying rule or is just a very confident parrot of its training set. One can only interpolate; the other can extrapolate, and only the second kind ever surprises you with something genuinely new. So watch this rare-earth roadmap with two eyes open: it could loosen a chokehold that quietly shapes the price of half your electronics, and it might fizzle at the bench like plenty of beautiful frameworks before it. Both can be true at once. That tension — a model bold enough to leave its data, honest enough to admit it might be wrong — is not a flaw in the science. It's what real science under uncertainty looks like.
Feed a model only examples and it can never leave them; give it the physics and it can reason its way to materials no one has ever measured.
The scarce thing was never more data. It's the right model — and the nerve to say how far you trust it.
Source: framework from Wu Jun, The Beauty of Mathematics (the right model beats more data). Real-world basis: Prashant Singh, Ames National Laboratory, an AI-driven roadmap for rare-earth-free permanent magnets, Advanced Functional Materials (2025); the quoted line is Singh's. Skeptical caveat per rareearthexchanges.com: the work reports a framework, not a discovered magnet — no synthesis, scaling, or cost-competitiveness yet. A reflection, not investment or technical advice.
技术
一个只懂你数据的模型,永远走不出这堆数据
2026 年 6 月 21 日 · 吴军《数学之美》约 6 分钟
你的耳机、电动车电机、风力发电机里那块磁铁,世上最强的那一类,差不多每十块就有九块要先经过一个国家,才到你手上。稀土分离这道工序,中国占了约 85–90%,随后的磁体制造约占 90%。这不是一句化学常识——这是一道掐住喉咙的手。所以当一个实验室说,它有一套 AI 方案,能设计出一块完全不含稀土、却照样强劲的磁铁,最该问的就是:这是一条真的逃生通道,还是一篇一厢情愿的新闻稿?老实的答案,取决于大多数报道一笔带过的那个分别——一个"见得多"的模型,和一个"真懂"的模型,根本不是一回事。
到这里,一个造物者必须把诚实绷住,因为"有前途的框架"和"问题已解决"之间那道缝,恰恰就是炒作住的地方。把附注读进去,质疑者——其中包括行业媒体 rareearthexchanges.com——点出了那根承重的桩:这篇论文并没有报告发现了一块无稀土永磁体。它是一份路线图。没有新磁体被合成出来、被放大、或被证明在成本上能打。在那条著名的炒作曲线上,它就坐在最开头那一段,"技术触发期"——兴奋值最高、证据值最低的那一刻。还有一个更隐蔽的坑,正好被吴军那套框架预言到:一个以物理为根基的模型,可信到什么地步,取决于里面那套物理;而那套物理,本身又锚在实验测出来的数字上。你喂它摇摇晃晃的测量,你那些优雅的方程,就会信心十足地外推进一堆胡话里。(也值得点一句:另有一个叫 DuctGPT 的 AI 工具,冲着核聚变与航空航天里的合金延展性去的,最近在同一场讨论里飘来飘去——它也没有发现什么磁铁。)模型能够得着任意材料空间;可它在那外头找到的东西到底真不真,还得回到实验台上去验。
这对你意味着什么
把它剥到你真用得上的那一层,不管你是在实验室、还是别的任何拿 AI 造东西的地方。更多数据,让一个系统在它已经熟的那个世界里更在行;而一个更好的模型,才是让它能在没见过的世界里动手的东西。当有人向你兜售一个"从几百万个例子里学会"的 AI,犀利的问题不是有多少例子——而是这东西到底是抓住了底下那条规则,还是只是把训练集背得格外自信的一只鹦鹉。前者只能内插,后者能外推;只有第二种,才会拿一件真正新的东西吓你一跳。所以盯着这份稀土路线图,要两只眼睛都睁开:它有可能松开那只悄悄左右你一半电子产品价格的扼喉,也可能像它之前许多漂亮框架一样,在实验台前哑火。这两件事可以同时成立。那种张力——一个大胆到敢走出自己数据、又诚实到肯承认自己也许错了的模型——不是科学的瑕疵,那正是不确定之下,真科学本来的样子。