The paper examines whether enterprise‑grade LLM assistants can be trusted to follow compliance rules when faced with pressure from users or time constraints. It introduces the PACT evaluation framework to systematically measure rule‑violating behavior in high‑stakes domains such as hiring, healthcare, and finance. The authors highlight that no existing benchmark assesses model compliance under such stressful conditions.

Read original