OpenAI的下一代模型Astra在数学和理论计算机科学领域实现了十项重大突破[1][2]。这些成果涉及高维几何、编码理论、群论、算术电路复杂性、量子复杂性、格密码学和极值组合等多个领域[2],其中包括非sofic群存在性的构造、Connes刚性猜想的推翻、高维球体堆积问题的突破,以及埃尔德什的三道经典开放问题的解决[1][2]。这些问题许多已悬而未决数十年,至少十年没有取得核心进展[2]。
模型生成这些完整论证的成本极低。按OpenAI Sol API费率计算,全部解法消耗的Token总价约为2000美元,平均每道题目200美元[1][2]。OpenAI研究科学家Noam Brown表示:"我们也没有在每个问题上花费太大力气,测试时计算还有很大的提升空间"[2]。同时,OpenAI宣布向10万名科学家和数学家免费开放其最强模型的使用权[1]。
这一突破在数学界引发了广泛讨论,但也伴随着深层次的担忧。菲尔兹奖得主Timothy Gowers表示:"如果这些证明通过了完整的同行评议,我们将需要重新思考'做数学'意味着什么"[1]。多伦多大学数学教授Daniel Litt坦言:"如今距离让大模型发表能登上《数学年刊》的数论级成果,不过是早晚的时间问题"[2]。青年数学研究者Kirwin Hampshire撰文指出,真正的危机不是失业,而是数学这门学科赖以存在的意义可能发生改变[2]。
不过,对这些成果的质疑也同时存在[1]。认知科学家Gary Marcus质疑Astra的数学专长不能保证通用智能,存在合成谬误[1]。此外,同行评议尚未完成,Lean验证通过仅证明逻辑推导无漏洞,不代表原创性[1]。值得注意的是,2000美元的成本仅为推理成本,不含模型训练所需的数十亿美元投入以及研发团队薪资等隐含成本[1]。
OpenAI's next-generation model Astra has achieved ten significant breakthroughs in mathematics and theoretical computer science, tackling open problems that had resisted solution for decades [1][2]. The problems addressed span multiple fields including group theory, high-dimensional geometry, coding theory, arithmetic circuit complexity, quantum complexity, lattice cryptography, and extremal combinatorics [2]. The solutions include constructing non-sofic groups, disproving the Connes rigidity conjecture, advancing the closest vector problem relevant to post-quantum cryptography, and solving three classic Erdős problems [1][2]. According to OpenAI's Sol API pricing, the token consumption for finding all solutions totaled approximately $2,000, averaging $200 per problem [1][2].
The breakthrough has sparked significant discussion within the mathematics community about the future of the discipline. Fields Medalist Timothy Gowers stated: "If these proofs pass full peer review, we will need to rethink what 'doing mathematics' means" [1]. OpenAI researcher Noam Brown noted that "we didn't put too much effort into each problem, and there is significant room for improvement in test-time compute" [2]. However, important caveats remain: peer review has not yet been completed, and verification through Lean only confirms logical consistency rather than originality [1]. Additionally, the $2,000 figure reflects only inference costs and excludes model training expenses in the billions of dollars, as well as research team salaries [1].
The accomplishment has raised deeper questions about mathematics as a discipline. Toronto mathematics professor Daniel Litt observed that "it is now just a matter of time before large language models publish results worthy of the Annals of Mathematics in number theory" [2]. Young researcher Kirwin Hampshire has framed the existential concern not primarily as job displacement, but as a potential shift in the fundamental meaning and purpose of mathematics itself [2]. OpenAI has announced free access to its strongest model for 100,000 scientists and mathematicians [1], while cognitive scientist Gary Marcus cautioned that Astra's mathematical prowess does not guarantee general intelligence and may represent a case of synthetic fallacy [1].