[Benchmarks] Qwen3.8-27B on one DGX Spark across SGLang, vLLM and llama.cpp
Article automatically generated from technical news.
I ran a 12-way Qwen3.8-27B comparison on one DGX Spark. Each engine used plain decoding plus MTP, DSpark, and DFlash2. The coding workload was a seeded 50-task HumanEval+ slice with thinking on, temperature 1.0, top-p 0.95, top-k 20, concurrency 1, and a 16,384-token completion ceiling. The numbers below are token-weighted net decode. Engine Plain MTP DSpark DFlash2 SGLang 12.65 24.20 27.96 **36.59** vLLM 10.95 20.89 25.58 **32.01*
Fonte originale