Benchmarks for the Intel Arc B70 demonstrate the performance of Qwen 3.6 35B (A3B) using INT4 Auto-round and MTP via vLLM with Intel XPU kernels. Testing shows throughput reaching 100 TPS at 4-10K prompt processing, scaling up to 120k context where token generation degrades to approximately 75 TPS. The setup utilized FP16 KV caching and specific quantization for router gates and shared experts.
Read original
reddit/r/LocalLLM