Z.ai has released GLM-5.3-Flash, a natively multimodal Mixture-of-Experts (MoE) model featuring 320B total parameters with 18B activated per token. The architecture employs a hybrid sparse-plus-linear attention stack and supports a 1M-token context window. This design allows for high-capacity performance while maintaining flash-tier serving economics.
Read original
reddit/r/machinelearningnews