GLM-5.3-Flash: How Z.ai Built a 320B MoE That Runs at 1/10th the Cost of Its Predecessor
Z.ai has released GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model with only 18 billion active parameters per token, achieving inference costs at roughly 1/10th of its predecessor. The model is released un…
→ View original source