llama.cpp b11344: CUDA Volta FA Bugfix Release
The b11344 release of llama.cpp addresses two broken cases when running using CUDA on NVIDIA Volta GPUs. No new features or breaking changes are introduced.
What changed?
llama.cpp b11344 resolves two previously broken function-asynchronous (FA) CUDA cases specifically affecting NVIDIA Volta GPUs. The release provides updated binaries for various platforms and hardware backends, including macOS, Linux, Android, and Windows. No new features, deprecations, or breaking changes accompany this update.
Why does it matter to an everyday developer?
If you are deploying AI workloads using llama.cpp with CUDA on Volta-series GPUs, this release resolves stability and correctness issues in two previously broken function-asynchronous cases. This ensures improved reliability for inference and experimentation on Volta hardware. For all other platforms and backends, functionality remains unchanged.
What can the developer do now?
If you are affected by the CUDA Volta FA cases, update to the b11344 release by downloading the appropriate binary for your target platform and backend from the official release page. No other migration or changes are required for those not using CUDA on Volta GPUs.
