自2017年Transformer架构问世以来,大语言模型的发展日益面临计算效率和能耗的挑战[1]。OpenAI今年计划在计算上花费50亿美元[1],而国际能源署预测数据中心总用电量将在2030年翻倍[1],这些因素促使多家初创公司开始探索超越Transformer的下一代技术方案[1]。
多个创新方向正在同时推进。Subquadratic开发的稀疏注意力机制模型SubQ声称是首个在搜索和编码等任务上与主流大语言模型相当的同类产品[1];Liquid AI的液态神经网络模型LFM由20%的Transformer和80%的液态神经网络组成[1],已获得近3400万次下载[1],可在价值50美元的树莓派上运行[1];Inception的扩散模型Mercury 2声称性能与OpenAI的GPT-4相当,但速度快10倍[1];Pathway基于状态空间模型开发的Dragon Hatchling在超过250000个数独谜题的基准测试中击败了97%的谜题[1]。这些初创公司通过各自独特的技术路径,试图在计算效率、上下文窗口和推理能力等方面突破Transformer架构的局限[1]。
A wave of artificial intelligence startups is pursuing alternatives to the Transformer architecture that has dominated large language model development since the 2017 publication of "Attention Is All You Need."[1] These companies are tackling fundamental limitations of current systems, including computational inefficiency and constrained context windows, through diverse technical approaches.
Among the emerging players, Subquadratic has developed a sparse attention mechanism called SubQ, which the company claims is the first to match mainstream LLM performance on tasks such as search and coding.[1] Liquid AI has taken a different path with its LFM model, which comprises 20 percent Transformer components and 80 percent liquid neural networks, and has achieved nearly 34 million downloads while remaining capable of running on a $50 Raspberry Pi.[1] Inception's Mercury 2 model purportedly delivers performance comparable to OpenAI's GPT-4 while operating ten times faster,[1] while Pathway's Dragon Hatchling defeated 97 percent of puzzles across a benchmark of over 250,000 Sudoku challenges.[1]
The push for architectural innovation reflects growing concerns about the resource demands of current systems. OpenAI plans to spend $5 billion on computing this year, according to company president Greg Brockman,[1] and the International Energy Agency projects that data center electricity consumption will double by 2030.[1] These startups represent competing bets on whether alternative architectures can deliver comparable capabilities with greater efficiency.