XHToken has released two new small language models, Spark-X2.5-4B and Spark-X2.5-1.7B, featuring a custom architecture rather than being fine-tunes of existing models. Both versions claim native 1M context window support, with the 4B model's benchmarks reportedly competitive with Qwen 3.5 9B. The models are currently available on Hugging Face but do not yet run out-of-the-box on llama.cpp, with support pending an open PR.
Read original
reddit/r/LocalLLaMA