Entropy-Driven Adaptive Router (EDAR): Hybrid Routing for RAG and Long-Context LLMs
EDAR is a framework that adaptively routes queries between retrieval-augmented generation (RAG) and full long-context (LC) processing in large language models. It leverages predictive entropy to balance accuracy and computational cost, providing a practical path to efficient LLM deployment for long-context tasks.
What Changed
The proposed Entropy-Driven Adaptive Router (EDAR) allows large language model deployments to choose at inference, for each query, whether to use fast retrieval-augmented generation (RAG) or escalate to more expensive long-context (LC) processing. It bases this routing decision on the predictive entropy—the uncertainty in the probability distribution—of the first few generated tokens from the RAG output.

Key Developer Takeaways
- EDAR requires only standard outputs from decoding, with no extra supervision or model retraining.
- The entropy threshold for routing is determined during setup using a held-out validation set, balancing cost with accuracy.
- EDAR is compatible with any retriever or LC model backbone—it is framework-agnostic.
- When used on LongBench v2 and Infinity-Bench, EDAR delivered 97.4% of LC accuracy at only 29.3% of the compute cost, escalating about 18.2% of queries for full LC processing.
- Predictive entropy (over initial RAG generations) correlates strongly with output hallucination risk. Routing based on this allows high efficiency without significant loss in answer accuracy.
