The team reduced the memory footprint of Qwen 3.5 4B to 800 MB while supporting a 128K token context, down from the typical 4 GB requirement. This was accomplished by compressing the KV cache and shipping a custom llama.cpp build that incorporates TurboQuant quantization. The changes were made as part of the Atomic Agent desktop release, which prioritizes local model execution over cloud‑first designs.
Read original
reddit/r/LocalLLM