The paper introduces Fairness Pruning, a lightweight structural intervention method for identifying and mitigating demographic bias in large language models. It employs minimally contrastive prompt pairs and inference-time activation capture to locate neurons that react differentially to demographic attributes within GLU architectures. The approach focuses on causal bias localization as an empirical foundation for future bias mitigation strategies.

Read original