The author trained a 3.87 B‑parameter Mixture‑of‑Experts model (1.45 B active parameters per token) from scratch on 86.5 B tokens, using a decoder‑only architecture with MoE in every layer, 32 layers, d_model = 2048, GQA (16 queries/4 keys/values), 16 experts and top‑4 routing, a 4096‑token context, and the Qwen3 tokenizer.

Read original