JARQ: Joint Alternating Refinement for Quantization Improves LLM Quantizer Accuracy
JARQ is a plug-in refinement method for group-wise post-training quantizers that improves accuracy and perplexity on quantized large language models, without additional inference cost.
JARQ (Joint Alternating Refinement for Quantization) introduces a refinement step for group-wise post-training quantizers applied to large language models (LLMs), such as Llama-2, Llama-3, and Qwen. Typical quantizers round weights onto a preset grid without realigning the grid to optimize for the resulting integer codes. JARQ starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with moves (via bounded Babai proposals) that adjust many codes of a group together.

This approach optimizes quantizer configuration without changing key deployment constraints: bit width, group structure, zero points, or inference cost. The refinement happens as a post-processing step and is compatible with major quantizers such as RTN, GPTQ, OmniQuant, and AWQ.
Key Results
- Accuracy and Perplexity
- JARQ lowered perplexity in 90 out of 96 experimental comparisons and raised mean multiple-choice accuracy in 23 of 24 settings across supported quantizers and models.
- Performance on 3-bit RTN
- Cut perplexity by up to 36%.
- Speed
- Improvement requires under a minute per 7B block.
- Compatibility
- Works with RTN, GPTQ, OmniQuant, AWQ, QEP, QuaRot, and OJBKQ.
