In a benchmark using a single RTX 3090, the Qwen 3.8 27B model (UD‑Q4_K_XL quantization) handled the heavy‑lifting of physics‑scene reasoning while Sonnet 5.5 via OpenRouter acted as the planner. This hybrid local‑cloud setup achieved the same task performance as a fully cloud‑based approach but at 2.7× lower cost. Measurements included llama.cpp built with CUDA, two parallel slots, 64K context, and turbo3 KV cache.

Read original