llama.cpp b11402: CUDA FlashAttention Efficiency Upgrade
The b11402 release of llama.cpp improves CUDA performance by preferring whole-tile FlashAttention scheduling for two-stage kernels, but disables certain builds for macOS (KleidiAI) and openEuler.
What changed?
llama.cpp release b11402 introduces an optimization for CUDA-based inference: it now prefers whole-tile FlashAttention scheduling for efficient two-stage kernels. This change is intended to increase performance for supported GPU setups. Additionally, the release disables macOS Apple Silicon (arm64, KleidiAI enabled) and openEuler builds.
Why does it matter to an everyday developer?
Developers leveraging llama.cpp with CUDA-enabled GPUs will benefit from improved inference speed due to more efficient FlashAttention scheduling. This is particularly relevant for those running large language models locally or in GPU-backed environments. However, users on macOS Apple Silicon with KleidiAI or on openEuler systems will not have a supported build in this release, which could limit or prevent upgrades on those platforms.
What can the developer do now?
If you use llama.cpp with CUDA, update to b11402 to potentially gain performance improvements. Prebuilt binaries for Windows and Linux CUDA versions are available in the release assets. Developers on macOS (non-KleidiAI) still have a supported build, but those relying on disabled builds will need to remain on earlier versions or await further updates. No code changes are needed to benefit from the CUDA kernel optimization.
