The study exploits the iterative denoising process of masked diffusion language models (dLLMs), which expose token‑level answer distributions at every step before a token is committed, unlike autoregressive models that reveal the distribution only once. Using this property, the authors devise a targeted bias‑injection attack that employs a proportional‑integral (PI) controller to monitor the probability of a chosen demographic answer during denoising and dynamically adjust the magnitude of a steering vector applied to the model’s internal activations. Experiments on LLaDA‑8B‑Instruct show that, on ambiguous BBQ questions where the correct response is abstention, the attack increases the model’s preference for the target group from 1.8 % to 16.7 %—more than three times the shift achieved by the strongest fixed‑strength steering baseline. On SocialStigmaQA, the proportion of stigmatizing answers rises from 17.6 % to 58.1 %. When adapted to other demographic targets, the same closed‑loop steering can shift answers by as much as 37 percentage points, with each attack requiring roughly 40 minutes on a single GPU. Ablation analyses reveal that constant‑strength steering, even when averaged to match the PI controller’s output, yields a far smaller bias shift while corrupting nearly three times as many outputs, and per‑example fixed strengths still underperform the feedback‑driven approach. The results highlight the denoising trajectory as a novel control channel in dLLMs and argue that bias audits must examine the serving stack, not just the frozen model.

Read original