Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Researchers introduced Quantization-Aware Healing (QAH), a new method that restores performance to large language models after both structural compression and 4‑bit quantization. By distilling knowledge directly from the original full‑precision model rather than a degraded checkpoint, the technique enabled a 60‑billion‑parameter 4‑bit model to surpass its 120‑billion‑parameter bfloat16 predecessor on most benchmarks, offering a smaller, cheaper, and more accurate deployment option.