The author complains that many new FP4‑optimized inference engines are advertised as fast but only work with NVFP4/MXFP4 formats, often resulting in hallucinated outputs. They argue that these niche solutions offer minimal real‑world benefit and ask the community to stop promoting them.
Read original
reddit/r/LocalLLaMA