The authors investigated whether training large language models to produce shorter chain‑of‑thought (CoT) rationales harms the transparency of those rationales. They fine‑tuned several pretrained models using three distinct efficiency‑inducing strategies: a fixed token budget applied uniformly across examples, a per‑example target length that varies with difficulty, and a group‑relative reward that encourages shorter CoTs relative to peers. After training, they measured CoT faithfulness by checking how well the generated rationale predicts the model’s answer on perturbed, related inputs, and assessed monitorability by determining whether the CoT still signals when an input intervention changes the final output. Results showed that faithfulness declined in most configurations, chiefly because the efficiency‑tuned models exhibited greater inconsistency in their reasoning across similar cases. In contrast, monitorability proved more resilient: even when the CoT was substantially compressed, the models continued to acknowledge the influence of input changes on their answers, preserving the ability to oversee behaviour. These findings suggest that length‑pressure training can undermine the reliability of CoTs as explanations while largely retaining their utility for detecting problematic input effects.

Read original