llama.cpp b11309: Hexagon ALLREDUCE Receives Safe Scatter Support and Optimizations
The b11309 release of llama.cpp optimizes the ALLREDUCE operation for Hexagon, introducing safe scatter mode support and code improvements to reduce register spills.
What's Changed
The b11309 release of llama.cpp introduces several targeted changes to the Hexagon backend's ALLREDUCE implementation:
- Added support for safe scatter mode in ALLREDUCE, enhancing operational robustness.
- Excessive comments in Hexagon ALLREDUCE code have been pared down for clarity.
- The implementation was rewritten to minimize register spills, improving efficiency.
Implications for Developers
These optimizations are relevant if you are running llama.cpp on Qualcomm Hexagon NPUs. The new safe scatter mode in ALLREDUCE may improve performance, stability, and debug clarity in distributed or multi-threaded scenarios involving this backend. There are no breaking changes.
Platforms and Availability
llama.cpp b11309 is available for macOS (Apple Silicon, Intel), iOS, Android, various Linux distributions and processor architectures (x64, arm64, s390x, Snapdragon), and Windows, with support for different acceleration backends including CUDA, Vulkan, ROCm, SYCL, and OpenVINO.
