reddit/r/LocalLLM

CPU inference DDR3/DDR4

u/Appropriate_Duck1778 2026-09-09

Wondering if anyone on here has actual benchmarks for CPU only inference DDR3 or DDR4 servers, im budget bound and my options are limited to legacy systems unfortunately. Heres what

→ View original source
reddit/r/LocalLLM

“The grim reality of corporate closed models: Gemini’s hardcoded filters completely break official legal research due to "Context Poisoning". Why local LLMs are the only way forward.”

u/Ill-Palpitation-9368 2026-09-04

A Reddit discussion highlights how corporate closed models like Google's Gemini suffer from "Context Poisoning," where hardcoded filters break legitimate professional workflows such as official legal research. The post a…

→ View original source
reddit/r/LocalLLM

VRAM goal reached... on a budget!

u/_TheWolfOfWalmart_ 2026-08-31

A Reddit user achieved 256 GB VRAM for $2,800 using 8x Radeon Pro V620 32 GB GPUs paired with dual Intel Xeon Gold 6148 CPUs, 384 GB DDR4 ECC memory, and a Supermicro X11DAi-N motherboard. However, the build suffers from…

→ View original source
reddit/r/LocalLLM

Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!

u/LLMFan46 2026-08-30

A community developer has released multiple uncensored GGUF-format models, including LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B and Qwen3.5-122B-A10B both with MTPs, and Qwen3-Coder-Next and Laguna-S2.1 wi…

→ View original source
reddit/r/LocalLLM

Qwen 3.8 27b harness

u/Dingydongy007 2026-08-30

The poster reports that multiple local deployments of the Qwen 3.8 27B model, including Opencode, qwen CLI, and Claude CLI, have yielded unsatisfactory coding performance, indicating difficulty achieving useful results. …

→ View original source
reddit/r/LocalLLM

Don't Sleep on EXL3 Quants

u/PyaesoneP 2026-08-30

A user reports high performance using EXL3 quants, specifically running the Muse Glimmer 30B model at 3.00bpw with a 100K context and Q8_O KV cache on a 12GB VRAM GPU. The setup achieves approximately 30 tokens per secon…

→ View original source
reddit/r/LocalLLM

Can I run Qwen 3.8 27b on this?

u/Hamonwrysangwich 2026-08-24

A user on the r/LocalLLM subreddit is inquiring about the hardware requirements necessary to run the Qwen 3.8 27b model. The post seeks community advice on whether their specific, though unspecified, system specification…

→ View original source
reddit/r/LocalLLM

Qwen3.8 27b - Holy crap!!

u/No-Manager1646 2026-08-16

The user wants a concise HTML summary of the Reddit post about Qwen3.8 27b. The post describes the user's experience running Qwen3.8 27b IQ4 NL at medium reasoning effort on their hardware (4060 8GB x4, 5060 Ti 16GB x8, …

→ View original source
reddit/r/LocalLLM

AMD Acquires Taalas

u/tcarambat 2026-08-06

AMD has announced the acquisition of Taalas, a startup that has raised $219 million in funding since 2023 and recently partnered with Cerebras at the Advancing AI Keynote. The purchase price has not been disclosed, but t…

→ View original source