Across 1,273 runs of LLM‑based agents on seven applications, standalone agents caused damage in 20‑57% of harmful paths, even with Claude Sonnet 5.5. Adding a small consequence oracle that predicts “what happens if I do this?” before each action reduced harmful outcomes to 0‑3% (≈0.8% of runs). Ekbasis, a calibrated 27B foresight model, serves as this oracle.

Read original