llama.cpp b11331: Enhanced CUDA Quantization and BF16 Support
The b11331 release of llama.cpp improves CUDA support for NVFP4 quantization, adds BF16 compute support for quantized models on compatible hardware, and adjusts precision handling for better performance on supported GPUs.
What changed?
- CUDA path now properly handles NVFP4 compute types, increasing compatibility and performance on NVIDIA GPUs.
- BF16 (bfloat16) compute type is automatically used for quantized models if the hardware supports it.
- For quantized models using NVFP4, accumulator precision is set to BF16 to match the minimum required range.
- Op_params preservation for per-expert matrix multiplication (matmul) is implemented for more robust computation flows.
- Pre-built binaries are now available for multiple platforms, but some openEuler builds remain disabled.
Sources
Why does it matter to an everyday developer?
CUDA users running quantized large language models like Llama will see improved accuracy, speed, and hardware utilization when running on GPUs that support the new compute types and BF16. The increased precision from BF16 is especially relevant to those deploying quantized models for inference, as it can help prevent numerical instability and maximize use of recent NVIDIA hardware. Developers using per-expert matmuls now benefit from better reliability thanks to parameter preservation.
