llama.cpp b11324 Reduces Memory Usage with Direct-IO Tensor Mapping
The llama.cpp b11324 release introduces a new memory optimization: it avoids creating a second full-size copy of each tensor by employing direct-IO for mmap, reducing RAM usage across supported platforms.
What changed?
The b11324 release of llama.cpp introduces an improvement to how tensors are loaded and stored in memory. The new approach uses direct-IO with mmap (memory mapping) to avoid keeping a second full-size copy of each tensor in RAM. This is intended to help reduce the memory footprint for applications using llama.cpp. The update is available for a wide range of platforms including macOS, Linux, Windows, Android, and iOS; openEuler builds are disabled for this release.
Why does it matter to an everyday developer?
For developers running or embedding llama.cpp models, especially on resource-constrained machines or when working with large language models, RAM usage is often a limiting factor. By preventing a second copy of each tensor from being loaded into memory, this update enables more efficient use of available system memory. This can let developers run larger models, serve more concurrent users per instance, or deploy llama.cpp in environments with lower hardware specifications.
What can the developer do now?
Developers can upgrade to llama.cpp b11324 to take advantage of reduced memory usage. Prebuilt binaries supporting this change are available for nearly all major platforms. To update, simply download the latest binary or rebuild from the updated source. There are no breaking changes or required code adjustments, but developers should note that openEuler builds are not included in this release.
