llama.cpp b11366: Improved Quantization Stability and New Regression Tests
The b11366 release of llama.cpp enhances quantization reliability by preventing invalid rounding in qkx3 scale search and adds regression coverage for degenerate imatrix cases. No breaking changes; binaries are updated for major platforms.
What changed?
llama.cpp b11366 introduces a key fix in the quantization module (ggml-quants): the scale search in qkx3 quantization no longer allows invalid (infinite, NaN, or out-of-range) values to be used for rounding. The quantization level is now clamped to a valid interval before rounding, avoiding assertion errors in debug builds and potential runtime instability. Additionally, this release adds regression test coverage for degenerate imatrix groups across several quantization types—including q2_K, q4_K, q5_K, q4_1, and q5_1—to improve correctness and catch edge cases early. Support for macOS Apple Silicon (KleidiAI enabled) and openEuler is marked as disabled, but binaries are available for a broad set of platforms.
Why does it matter to an everyday developer?
This update significantly increases the reliability and predictability of model quantization steps in ggml-based LLM deployments. Developers running llama.cpp with custom or edge-case inputs are less likely to encounter quantization failures or assertion-triggered crashes in debug scenarios. The improved test coverage strengthens confidence that rare input patterns will not produce unexpected quantization behavior. There are no breaking changes, so existing workflows and integrations will continue to work as expected.
What can the developer do now?
Developers are encouraged to update their llama.cpp deployment to b11366 to benefit from more stable quantization and broader regression test coverage. Existing binary packages are available for all major platforms except those newly disabled (macOS Apple Silicon with KleidiAI and openEuler). No special migration steps are needed. If your workflow involves quantizing models or running debug builds, this release reduces the likelihood of hitting previously unhandled quantization edge cases.
