字节跳动 Seed 团队在 arXiv 发表研究论文,揭示了 DeepSeek 模型采用的分块 KV 缓存压缩技术存在的性能缺陷1。该技术通过固定步幅将连续标记窗口压缩为更少缓存条目,但研究发现这一方案存在相位敏感性问题,导致长上下文检索性能出现周期性变化1。
研究评估了 DeepSeek-V4-Flash、DeepSeek-V4-Pro 和 DeepSeek-V4.1-Flash 模型,发现模型中存在系统性不对称性,相同信息在某个阶段容易被检索,但在另一阶段则难以被检索1。论文显示,长上下文检索准确度在各阶段之间的差异最高可达 40 个百分点1。这一发现表明,仅依赖平均基准分数可能会掩盖模型在实际应用中的性能弱点1。
ByteDance's Seed team has published research on arXiv revealing a critical vulnerability in DeepSeek's architecture.1 The researchers evaluated DeepSeek-V4-Flash, DeepSeek-V4-Pro, and DeepSeek-V4.1-Flash models and discovered that the chunked KV cache compression technique employed by these models exhibits phase sensitivity issues.1 This technique works by compressing consecutive token windows into fewer cache entries using fixed strides, but the researchers found it introduces systematic asymmetries in the models' behavior.1
The phase sensitivity problem manifests as periodic performance degradation in long-context retrieval tasks.1 The same information can be retrieved easily at one stage but becomes significantly harder to access at another stage, with retrieval accuracy differences reaching as high as 40 percentage points across different phases.1 These findings suggest that average benchmark scores commonly used to evaluate model performance may mask underlying performance weaknesses that emerge in real-world long-context scenarios.1 The research was submitted to arXiv at the end of September 2024.1
评论
还没有评论,欢迎留下第一条。