The Kimi K3 full model was successfully deployed on a 16x GB10 cluster using dspark, achieving an average speed of 20+ tps and a peak of 38 tps. The setup also demonstrated a prefill rate of 750 tps. A vLLM image and implementation instructions are expected to be released following further optimization tests.
Read original
reddit/r/LocalLLaMA