The paper introduces Decodability Supervision, a method using natural-language autoencoders to evaluate activation explanations by reconstruction fidelity. However, it reveals the test is structurally insensitive to false claims:
huggingface/daily-papers