Debias-SparseGPT is a new post-training pruning method designed to mitigate the amplification of biases often caused by weight sparsification techniques like SparseGPT. The approach incorporates representational debiasing through a second-order term to ensure model outputs remain consistent regardless of persona cues in prompts.
Read original
huggingface/daily-papers