CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies Researchers introduce CoffeeBench, a novel benchmarking framework designed to evaluate the performance of Large Language Model (LLM…
→ View original source