The author deployed DeepSeek V4 Flash Vision across four unlocked NVIDIA CMP 170HX 64GB GPUs, providing 256 GB of aggregate HBM2e memory. Using vLLM with pipeline parallelism = 4, a 262 K token context, FP8 KV cache and DSpark speculative decoding (6‑token lookahead), the system achieved ~50–70 tokens/s on typical coding tasks and peaks near ~95 tokens/s. The setup runs on Ubuntu 24.04 with OMP as the coding‑agent frontend, limiting concurrent sequences to two.

Read original