研究人员对混合专家(MoE)语言模型进行了深度分析,发现当被问及代码是否恶意时,这些模型会激活与道德判断相关的神经路径1。Manifold研究员Cody Nash领导的研究团队通过路由移植实验和工作空间透镜分析,在三个开源模型——OLMoE、DeepSeek-V2-Lite通用版和编码版上测试了这一现象1。研究基于240个代码样本进行,包括恶意和良性PyPI包、易受攻击和修复后的函数等多种情景1。
研究结果表明,恶意问题的路由距离道德问题最近,在每个模型上都是1.0的同义词单位距离1。通过移植其他问题的路由,研究人员验证了模型确实调用了道德推理机制而非仅进行代码语法分析,尽管移植操作会改变模型答案,但不会传递供体问题的答案或概念1。专家剪枝实验进一步确认了道德路径的关键作用——REAP剪枝删除了524个道德路径细胞中的116个,模型的判断能力仍保持完美(AUROC为1.00),而GPT-OSS-20B更激进的剪枝在保留路径专家的同时导致判断能力显著下降(AUROC从1.00降至0.71),这表明道德路径是路由器的关键咨询机制1。
Researchers studying mixture-of-experts (MoE) language models have discovered that when asked whether code is malicious, these systems engage neural pathways associated with moral judgment rather than relying solely on syntactic code analysis.1 The finding comes from work conducted at Manifold by researcher Cody Nash and team members who employed routing transplantation experiments and workspace lens analysis to trace model behavior.1
The study tested three open-source models—OLMoE, DeepSeek-V2-Lite in both general and coding variants—across 240 code samples that included both malicious and benign PyPI packages as well as vulnerable and patched functions.1 Routing distance analysis revealed that queries about malicious code clustered nearest to moral reasoning questions, consistently achieving a synonymy distance of 1.0 across all tested models.1 When researchers transplanted routing patterns from unrelated queries, the model's answers changed, though the transplanted routing did not transfer the donor question's answers or conceptual framework.1
Targeted pruning experiments further confirmed the importance of moral pathways in the decision process.1 When REAP pruning removed 116 of 524 moral pathway cells, the model maintained perfect discrimination between malicious and benign code with an area under the receiver operating characteristic curve of 1.00.1 However, a more aggressive pruning strategy on GPT-OSS-20B that preserved pathway experts resulted in a significant degradation of judgment capability, with the score dropping from 1.00 to 0.71.1
评论
还没有评论,欢迎留下第一条。