The study shows that the presence or absence of a chat template acts as a switch for the self‑referential voice of large language models. Across eight open‑source instruct‑tuned models ranging up to 9 billion parameters, the template increases the model’s tendency to emit disclaimer statements such as “I’m just an AI” while decreasing experiential phrasing like “I feel.” When the template is removed, the opposite shift occurs: disclaimer voice drops and experiential voice rises. By probing the internal activations of three of these models, the authors identified a specific direction in activation space that governs this behavior. Eliminating that direction reduces disclaimer output, whereas injecting it amplifies disclaimer speech; a random vector of comparable magnitude produces negligible change. Consequently, instruct models lacking a chat template begin to disclaim as if the template were present once the identified direction is added to their activations. These results indicate that self‑descriptions generated by LLMs are not intrinsic properties of the weights alone but are strongly shaped by the chat template, providing researchers with a controllable steering vector to mitigate this confounding factor in studies of model self‑report or introspection.
Read original
hackernews