开发者发布了PSSA,一个用Rust从零编写的小型非Transformer语言模型1。该模型采用递归状态空间层读取文本,配备可查询的情节记忆库,并能在运行时重写部分权重1。模型参数规模为1.5M,目前处于研究原型阶段1。
PSSA在性能表现上明显优于同等规模的Transformer模型1。在12.7M token的WikiText-103语料库上,PSSA的训练交叉熵为3.98,而Transformer为4.43,相差0.45 nats;困惑度方面,PSSA为53.7,Transformer为83.71。在生成效率上,PSSA在相同CPU上生成200个token的速度约为Transformer的12倍,在Kaggle T4 GPU上的训练速度约900 tokens/秒1。
该项目完全用Rust手写线性代数,不依赖任何机器学习框架1。模型整合了递归状态空间核心、情节记忆库、可塑权重和闭式巩固等特性1。
A developer has released PSSA, a compact language model designed as an alternative to transformer architectures and implemented entirely in Rust without relying on machine learning frameworks 1. The model employs a recurrent state-space core, an episodic memory system, and trainable weights that can be modified during inference 1.
Performance benchmarks demonstrate significant advantages over comparable transformer models. When trained on the WikiText-103 corpus containing 12.7 million tokens, PSSA achieved a cross-entropy loss of 3.98 compared to a transformer's 4.43, a difference of 0.45 nats 1. The perplexity metric shows even starker separation, with PSSA reaching 53.7 against the transformer's 83.7 1. On CPU hardware, PSSA generates text approximately 12 times faster than transformer models, while maintaining a parameter count of 1.5 million 1. Training speed on Kaggle T4 GPU infrastructure reaches roughly 900 tokens per second 1.
The implementation features custom linear algebra routines written entirely in Rust, incorporating architectural innovations including a recurrent state-space foundation, an episodic memory bank for queryable recall, plasticity mechanisms enabling weight adaptation at runtime, and closed-form consolidation techniques 1. The project remains in the research prototype phase 1.
评论
还没有评论,欢迎留下第一条。