MiniMax M3 Prompt Caching Now Available on SambaCloud
SambaCloud enables prompt caching for MiniMax M3, improving inference performance and reducing costs for requests with shared prompt prefixes of at least 4,096 tokens and supporting context windows up to 192k tokens with zero code changes required.
What changed?
Prompt caching is now available for the MiniMax M3 model on SambaCloud. When multiple requests share a stable prompt prefix of at least 4,096 tokens, SambaCloud serves the cached prefix rather than recomputing it. This feature requires no changes to application code. Cached tokens are billed at a rate 90% lower than the standard input token rate. This capability works across context windows from 8k up to 192k tokens.

Why does it matter to an everyday developer?
Developers building applications with repeated or large prompts will see reduced time to first token (TTFT) by 35% to 88%, resulting in faster response times and improved user experience. Cost is also significantly reduced for cached tokens, cutting inference expenses on workloads with repeated context. No application changes are required to benefit from these improvements, simplifying integration.
What can the developer do now?
If you deploy MiniMax M3 on SambaCloud, your workloads automatically benefit from prompt caching where prompt prefixes of at least 4,096 tokens are repeated across requests. Review your usage and context window sizes to optimize cost savings and latency gains. No code modifications are needed to activate or use this feature.
MiniMax M3 Prompt Caching on SambaCloud
| Feature | Before | After |
|---|---|---|
