llama.cpp b11330: Qwen4Exp Adds MTP Support and Internal Improvements
The latest llama.cpp release (b11330) introduces MTP support for Qwen4Exp, streamlines state management, and cleans up code, benefitting developers using this open-source LLM runtime.
What changed?
llama.cpp b11330 introduces MTP (Most-Token Prediction) support for the Qwen4Exp model, removes the redundant has_state member (using ctx_bufs presence instead), standardizes naming, and performs code clean-up by improving comments and recurrent memory handling.
Why does it matter to an everyday developer?
These updates mean smoother and potentially more efficient experience when running Qwen4Exp models, especially if your workloads can leverage MTP. The internal clean-up reduces code complexity and potential bugs for those customizing or contributing to llama.cpp. State detection is now less error-prone, relying on actual context buffers instead of a flag, which lowers state sync issues.
What can the developer do now?
- Upgrade to llama.cpp b11330 to benefit from MTP in Qwen4Exp and internal improvements.
- If developing against Qwen4Exp, test MTP-enabled workflows and monitor for improved inference behaviors.
- For contributors, adapt to the new state logic (using ctx_bufs) and follow the updated naming conventions.
