A new pull request (#28770) for ggml-org/llama.cpp introduces CUDA‑enabled sparse flash attention for the Qwen4 model, aiming to improve inference speed. The change was highlighted by u/jacek2023 on the r/LocalLLaMA subreddit as another Qwen Flash Next speedup.

Read original