Kimi K3 Architecture: Scaling 2.8T Parameters for Long-Context Multimodal MoE

Article automatically generated from technical news.

Introduction: Scaling Along Length, Depth, and Width Scaling frontier language models requires balancing parameter capacity, context window length, and computational efficiency during training and inference [1]. Kimi K3 by Moonshot AI is a native multimodal Mixture-of-Experts (MoE) model with 2.8 trillion total parameters, 104 billion activated parameters per token, and a 1-million-token context w

Fonte originale