Moonshot AI has launched Kimi K3, a 2.8‑trillion‑parameter open MoE model that it claims is the world’s first 3‑trillion‑parameter open model. The system incorporates Kimi Delta Attention, a hybrid linear attention mechanism that bypasses conventional prefix caching, and has been integrated into vLLM for up to 6.3× faster decoding at its 1‑million‑token context window. Kimi K3 also features depth‑wise attention residuals that selectively retrieve representations across layers.
Read original
reddit/r/machinelearningnews