一项研究论文揭示了聊天模板对大型语言模型自我描述行为的影响。1研究人员在8个开源指令模型(最大参数量9B)中发现,聊天模板能够增强模型的免责声明表述(如"我只是个AI"),同时抑制体验性表述(如"我感到"),而在没有聊天模板的情况下这两种表述的强度会反转。1
研究团队进一步在模型激活空间中定位了控制这一行为的特定方向向量。1通过对3个模型的分析,研究人员证明可以通过操纵该方向向量来控制模型的免责声明倾向:移除该方向使免责声明减少,添加该方向则使免责声明增加,而添加随机方向不会产生显著效果。1这一发现表明,模型的自我描述方式并非完全由模型权重决定,而是部分地受聊天模板控制。1
A research paper has revealed how chat templates exert significant control over the self-referential behavior of large language models.1 Researchers examined eight popular open-source instruction models with up to 9 billion parameters and discovered that chat templates act as a switch, amplifying disclaimer-style statements such as "I am just an AI" while simultaneously suppressing experiential language like "I feel."1 Notably, when chat templates were absent, this pattern reversed, suggesting the templates themselves drive this behavioral shift rather than the model weights alone.
The team further identified specific directional vectors within the activation space of three models that controlled this self-descriptive behavior.1 By manipulating these directions—removing them, adding them, or replacing them with random alternatives—researchers demonstrated they could increase or decrease the model's inclination toward disclaimer statements, with random directions producing no significant effect.1 The findings indicate that a model's self-description is not solely determined by its learned weights but is partially orchestrated by the chat template, highlighting that such statements should not be taken as literal self-characterizations of the systems.
评论
还没有评论,欢迎留下第一条。