KronQ is a new post-training quantization (PTQ) framework for large language models that moves beyond traditional second-order methods like GPTQ. By incorporating gradient covariance into the quantization pipeline, the method addresses the limitation of assuming uniform output channel contribution to the layer-wise reconstruction objective. This approach utilizes Kronecker-factored Hessian information to improve compression accuracy.
Read original
huggingface/daily-papers