XingChen-AGI released Xing4.0-29B-A4B, a Mixture-of-Experts large language model with 29 billion total parameters and 4 billion activated per token, supporting a native 256 K token context length extendable to 512 K. It is the first model of this scale trained entirely on Huawei’s Ascend NPU platform using the MindSpore framework.

Read original