This paper introduces ToolHazard, a scalable adversarial environment synthesis framework designed to evaluate the security and alignment of LLM-based agents integrated with external tools. Unlike prior approaches that rely on manual environment construction or stochastic LLM-based tool simulation, ToolHazard automates the generation of diverse adversarial scenarios with dynamically placed prompt injections. The framework aims to enable broader and more systematic security research across varied domains by reducing human engineering overhead and expanding beyond predefined injection locations.
Read original
huggingface/daily-papers