llama.cpp b11374: Major OpenVINO Backend Update with Performance and Profiling Enhancements
llama.cpp b11374 upgrades the ggml-openvino backend to OpenVINO 2026.4.1, adds inference profiling, expands supported operations and device handling, improves MoE model performance, and introduces breaking changes in device error handling.
What changed?
The b11374 release of llama.cpp substantially updates its OpenVINO backend (ggml-openvino): - ggml-openvino is now based on OpenVINO 2026.4.1, with a new graph cache, broader op coverage, and optimized device enumeration and memory reporting. - Inference profiling tools are introduced for detailed performance analysis. - Remote output tensors are used by default; single-sequence recurrent state handling is optimized. - MoE (Mixture-of-Experts) routing and fusion for models with fused gate-up weights (notably improving Qwen3.5 and Gemma model performance). - New quantization and convolution ops, improved handling of HARDSIGMOID and EXPM1. - Breaking device selection behavior: specifying an unavailable GGML_OPENVINO_DEVICE now gives a clear error listing available devices (no more silent CPU fallback). - Only explicitly selected OpenVINO devices will be used for GPU offload, with others reported as IGPU. - 'cache_only' mode for loading models directly from compiled caches. - More robust device listing, logging, and driver/version updates. Known limitations remain around specific MoE fusion cases and correctness for some model configurations (e.g., Gemma 26B).
Why does it matter to an everyday developer?
The update improves llama.cpp’s usefulness for local inference on Intel and compatible hardware: - Stronger MoE model performance (notably for Qwen3.5, Gemma) with fused ops—up to 24× throughput gains in certain configurations. - Detailed inference profiling means developers can troubleshoot and optimize models effectively. - Expanded supported ops and improved memory/device reporting increase backend reliability on diverse systems. - Tighter device selection and error reporting reduces "silent misconfiguration" risks, ensuring models run where intended. However, code or workflows relying on fallback behaviors may break and require explicit configuration. - New capabilities like 'cache_only' mode streamline deployments where loading speed is critical.
