The paper demonstrates that LLM safeguards based solely on copyable context cannot reliably prevent misuse, because the same model output can aid both authorized users and attackers. This creates a safety trilemma where useful capability, reliable safety, and open access cannot coexist. Adding hard‑to‑copy trusted credentials that predict actual downstream use can strengthen safeguards and eliminate the worst‑case assistance floor for attackers.

Read original