Open sourcellama.cpp
llama.cpp b11306 Released: Improved Graph Shape Testing in test-llama-archs, openEuler Support Disabled
llama.cpp b11306 introduces enhanced testing for graph shape changes in test-llama-archs and disables openEuler builds for several configurations. Developers should review platform support and consider the updated graph shape validation mechanism.
Key Changes
- test-llama-archs now toggles the causal_attn flag after device decode to help catch graph shape changes.
- The script decodes n_ubatch/2, then n_ubatch tokens (with equal node counts), detecting unwanted graph shape reallocations under GGML_SCHED_NO_REALLOC.
- This behavior is skipped for encode architectures.
- openEuler builds are disabled for all x86 and aarch64 configurations.
Platform Support
Binaries for llama.cpp b11306 are available for macOS (Apple Silicon arm64, Intel x64, iOS), Linux (Ubuntu x64/arm64/s390x, Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, SYCL FP32/FP16, Snapdragon), Android (arm64, Snapdragon), Windows (x64/arm64 CPU, CUDA 12/13, Vulkan, OpenCL Adreno, ROCm 10.0, OpenVINO, SYCL), and a UI package. All openEuler builds are now disabled.
Impact for Developers
- Graph shape consistency issues in test-llama-archs will now be easier to catch for architectures using device decode.
- Developers targeting openEuler must use alternative supported platforms.
- Existing test/CI infrastructure should account for the enhanced graph shape testing to minimize false negatives.
