Existing bias audits based on model outputs can be expensive and may miss internal representation changes. The paper proposes a reference-based method that compares bias in hidden-state representations across related model variants by encoding sentences according to their similarities to a reference, addressing representation-geometry shifts caused by fine-tuning. Read original