DeepSeek is training a 2‑trillion‑parameter model and plans to eventually build an 8‑trillion‑parameter model. Its current offerings include a 552‑billion‑parameter Flash model and a 1.6‑trillion‑parameter Pro model with 49 billion activated weights per token, while a Mythos/Fable variant is estimated at around 10 trillion parameters.

Read original