Weave Router 2.0 is an open‑source model‑routing system that plugs into coding agents such as Claude Code or Codex and dynamically selects among a pool of LLMs—using Astra for demanding debugging or system‑design tasks and Deepseek v4 Flash for routine frontend edits—to achieve Astra‑level performance while reducing cost and latency. Benchmarked on Terminal Bench 4.0 and SWE Atlas against GPT‑6 Astra, the router matched Astra’s pass‑rate on both suites, consuming only 52 % of Astra’s cost on Terminal Bench (2.2× speedup) and 54 % of its cost on SWE Atlas (2.5× speedup). Performance gains stem from three key upgrades: (1) a new architecture that first trains a hidden Markov model to capture session state trajectories, then feeds those states to a classifier that clusters similar model choices, drastically shrinking the routing search space from ~10^100 possible paths to a tractable bucket selection problem; (2) an expanded training set generated by frontier LLMs to label diverse coding‑agent sessions, providing richer supervision for both the HMM‑classifier and reinforcement‑learning components; and (3) a refined cache‑eviction impact module that accurately estimates the expected value of model switches, minimizing unnecessary cache fills while still triggering switches when beneficial. The router’s code is available at https://github.com/weave-os/router, with a hosted demo at https://weaveos.com/router.
Read original
hackernews