The llama.cpp PR #26136 adds native lossless 2.05× compression for FP32 weights, reducing them to 15.64 bits per weight while preserving exact bit‑for‑bit reconstruction and baseline perplexity. It also achieves 1.14× lossless compression for BF16 models, down to 14.07 bpw. This enables unquantized model storage and inference without quality loss.

Read original