DiSCO (Defending text-to-image generation through distribution-guided contrastive prompt optimization) offers a black-box defense that mitigates NSFW content and red‑teaming attacks in text‑to‑image models by optimizing prompts with distribution‑guided contrastive learning. Unlike prior white‑box methods that rely on text encoder optimization, weight editing, or inference‑time interventions, DiSCO works without internal model access, enabling scalability to proprietary systems.
Read original
huggingface/daily-papers