Aleph Alpha’s Kolibri is an open‑weight mixture‑of‑experts large language model released on 3 October 2026 under the Apache 2.0 license. It contains 78.1 billion parameters organized into 50 layers, each with 384 expert sub‑networks plus one shared expert; a router selects six experts plus the shared expert for every token, yielding roughly 3.46 billion active parameters (≈4.4 % of the total) per token. The model was trained from scratch on about 24 trillion tokens, roughly one‑fifth German, using 768 NVIDIA B200 GPUs in Germany and Finland. Kolibri’s UniBPE tokenizer (128 k vocabulary) reduces German token count by 11.2 % compared with GPT‑5’s o200k_base tokenizer. Forty of its layers employ sliding‑window attention over a 512‑token window, with every fifth layer using full attention, enabling a native context of 262 144 tokens and validated performance up to 1 048 576 tokens. It supports tool calling, four reasoning effort levels (none‑high), and a Merlin‑Arthur protocol that yields a 44 % abstention rate on the Omniscience test. Evaluation shows overall English score 75.5, German 70.8, AIME 2025 English 96.9 and German 87.5, and a 1 M‑token RULER base‑model score of 63.2. The model requires ≈78 GB of FP8 weights (e.g., two 80 GB A100/H100 GPUs or a single H200/B200/B300) and is served via Aleph Alpha’s vLLM plugin.
Read original
hackernews