The author conducted multi‑hour llama.cpp optimization experiments on Qwen3.6‑35B‑A3B and Qwen3.8 Flash‑Next MoE models, achieving substantial speedups in prompt processing and source‑code editing while keeping model weights and quantization unchanged, though some regressions and tradeoffs were observed. Tests were run on an RTX 4080 (16 GB VRAM), Ryzen 9 5900X, and 64 GB DDR4 RAM, with source patches, benchmarks, and reproduction guides shared publicly.
Read original
reddit/r/LocalLLM