GLM-5.3-Flash (320B) runs on an M5 Ultra Mac Studio via llama.cpp, achieving ~61 token/s generation with its MTP head (probabilistic drafting, 3 draft tokens), up from ~39.7 token/s without it. Prompt processing is slower, at ~765 token/s for 4K tokens and ~630 token/s for 59K tokens, with a 29K-token prompt taking ~45 s, while the model accommodates the full 262K‑token context with ~15 GiB memory to spare and maintains correct recall up to 58.8K tokens. Limitations include the inability to enable image input and MTP simultaneously, and challenges with deep‑context performance on Metal.

Read original