Can ChatGPT win a Fields Medal? | ChatGPT 能否获得菲尔兹奖? - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
FT英语电台

Can ChatGPT win a Fields Medal?
ChatGPT 能否获得菲尔兹奖?

New AI models could soon pose a threat to the world’s top mathematicians
新型人工智能模型在解决数学难题方面的能力正迅速提升。其在极具挑战性的全新问题上的表现,已对全球顶尖数学家构成威胁。
00:00

undefined

The writer is a science commentator

When Yang-Hui He, a fellow at the London Institute for Mathematical Sciences, received an invitation to an all-expenses paid weekend held in Berkeley, California, last month, it was a no-brainer. The trip would afford the Oxford university lecturer, an expert in algebraic geometry and string theory, insider access to a potentially historic moment for his discipline.

Plus, the brief sounded fun: working with other top mathematicians to find out if the most advanced AI models, when confronted with brand new problems, could rival or exceed the collaborative reasoning abilities of the best human minds. The answer? The machines did better than expected. “I’m not saying we felt existentially threatened but there was a general feeling of awe,” He told me. He also flew back $1,500 richer after dreaming up a problem that stumped the AI. 

Using AI to crack maths puzzles is not new. In early 2024, Google DeepMind unveiled technology that could hold its own in high-school student maths competitions. But interacting with the latest AI models last month felt more “like working with a very, very good graduate student”.

This moment could potentially change the profession. While the prospect of a machine securing a Fields Medal — widely regarded as mathematics’ equivalent of a Nobel Prize — still feels reassuringly distant, one can envision an unsettling future in which graduate maths programmes are pruned, university departments are shuttered, and the torch of Pythagoras and Euclid passed to a faceless silicon successor.

The weekend in mid-May was organised by Epoch AI, a US-based non-profit organisation that benchmarks AI capabilities. In an initiative set up last autumn called FrontierMath, Epoch paid professional mathematicians to submit novel problems along with their solutions, proofs and derivations, that could be used to challenge AI models.

These specially crafted conundrums, earning their creators up to $1,000 apiece and graded into three tiers of difficulty (including undergraduate and research level), were collected by Epoch via the secure messaging app Signal, so that they could not be inadvertently included in AI training data scraped from the internet. By April this year, Scientific American reported, an OpenAI model had confounded expectations by solving around a fifth of them.

And so it was time for tier 4 challenges: super-tricky problems that would take top academics weeks or months to solve collaboratively — and designed to resist AI guesswork or brute force number-crunching. Thirty academic experts, including He, met at Epoch’s Berkeley offices to brainstorm some new problems in person. Again, secrecy prevailed: lunches and dinners were brought in; attendees signed non-disclosure agreements and He recalled needing security cards to visit the toilets.

The full results of how the AI model performed on 50 tier 4 problems are yet to be disclosed. But He was struck by how much the tech has improved since 2022, when “ChatGPT couldn’t even find the tenth digit of seven divided by 13 . . . now it’s beginning to do something more intelligent.”

He explained how the AI, called o4-mini, was able to solve some of the problems in minutes, writing mathematical scripts and drawing on external specialist software. Most impressive, he said, were detailed literature searches, turning up obscure but critical papers and coding shortcuts. Another attendee, Ken Ono, a University of Virginia mathematician and freelance consultant for Epoch, called the results “frightening”. 

The project is not without controversy: in January, Epoch apologised for initially failing to disclose OpenAI’s financial backing of FrontierMath, leading to suspicions that the company’s AI models, including o4-mini, would have favoured access to some of the unseen maths problems used for benchmarking.

AI models cannot yet tackle the hardest maths challenges. Even so, one can imagine the next generation of machines thinning out the next generation of human mathematicians. That could shrink the pool from which future Fields medallists are drawn; there might be fewer hopefuls to attack famous unsolved problems like the Riemann Hypothesis, one of six carrying a $1mn bounty.

While the use of prime numbers in encryption shows the practical use of mathematics, there is something quite profound about living in a universe filled with dazzling concepts like zero, infinity and imaginary numbers. Perhaps fretting over whether the addition of AI might subtract from this human endeavour is not that irrational after all.

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

印度的俄罗斯石油难题

面对特朗普的施压,莫迪要么接受美国的关税,要么从俄罗斯转向其他供应国,要么尝试与特朗普达成某种妥协。

梅德韦杰夫:从自由派希望到核威胁的漫长转变

这位俄罗斯前总统在社交媒体上公开挑衅特朗普。他已从俄罗斯自由派和白宫决策者眼中的白衣骑士,变成了莫斯科的“核战争狂人”。

数千公司董事在工党税改后离开英国

FT的分析发现,从去年10月到上月,有3790名公司董事报告离开英国,阿联酋成为首选目的地。

局长遭解雇前,美国劳工统计局就已然陷入困境

经费和人员削减阻碍了美国劳工统计局汇总重要报告的能力。

FT社评:特朗普对经济数据的攻击产生寒蝉效应

特朗普解雇美国劳工统计局局长会破坏关键国家统计数据的公信力。

特朗普的阿拉斯加州液化天然气项目未能打动亚洲盟友

日本和韩国抵制美国压力,拒绝将对管道项目的承诺纳入贸易协议。
设置字号×
最小
较小
默认
较大
最大
分享×