llama.cpp b11387 Fixes n-gram Draft Rejection at Higher Temperatures
Release b11387 of llama.cpp includes a bug fix for n-gram drafts being incorrectly rejected at temperatures greater than zero after truncation. This brings more reliable outputs for developers using sampling at non-zero temperatures.
What changed?
With version b11387, llama.cpp addresses a bug where n-gram drafts were being rejected when running with temperature (temp) greater than zero after truncation. The fix ensures that n-gram-based sampling works as intended at higher temperatures, rather than failing or producing unintended results in these scenarios.
Why does it matter to an everyday developer?
If you use llama.cpp for local inference and set sampling temperature above zero to introduce output diversity, this bug could have resulted in rejected draft generations, reducing output quality or consistency. The update makes sampling more predictable and reliable, particularly when using creative or varied outputs.
What can the developer do now?
Upgrade to llama.cpp b11387 or later to ensure your n-gram sampling logic functions correctly at all temperature settings. You can download updated binaries for all supported platforms, including macOS, Linux, Windows, Android, and iOS, from the release page.
