Running Qwen 3.8 27B FP8 via vLLM (with prefix caching and speculative decoding) on a single RTX PRO 6000 Blackwell 96GB, a small developer team processed about 2 billion tokens in the first week, primarily for agentic coding workloads. Two individual users contributed roughly 950 million and 900 million tokens respectively. The workflow involves agents loading repositories, invoking tools, editing code, running tests, and repeatedly re‑sending context.
Read original
reddit/r/LocalLLM