Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!

Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!

u/LLMFan46 2026-08-30

A community developer has released multiple uncensored GGUF-format models, including LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B and Qwen3.5-122B-A10B both with MTPs, and Qwen3-Coder-Next and Laguna-S2.1 wi…

→ View original source

Qwen 3.8 27b harness

u/Dingydongy007 2026-08-30

The poster reports that multiple local deployments of the Qwen 3.8 27B model, including Opencode, qwen CLI, and Claude CLI, have yielded unsatisfactory coding performance, indicating difficulty achieving useful results. …

→ View original source
Don't Sleep on EXL3 Quants

Don't Sleep on EXL3 Quants

u/PyaesoneP 2026-08-30

A user reports high performance using EXL3 quants, specifically running the Muse Glimmer 30B model at 3.00bpw with a 100K context and Q8_O KV cache on a 12GB VRAM GPU. The setup achieves approximately 30 tokens per secon…

→ View original source
Loading more articles...