Less Uniform Discrete Diffusion (LUDI) Improves Diffusion Language Model Scalability and Speed
A new framework called Less Uniform Diffusion (LUDI) addresses key limitations in uniform diffusion language models (UDLMs) by introducing a less uniform loss and per-token time embeddings, resulting in faster and more effective language generation.
The Less Uniform Diffusion (LUDI) framework introduces important changes to the architecture and training of Uniform Diffusion Language Models (UDLMs). Previous UDLMs struggled to scale due to an overly uniform training objective and confusion between condition and target during sampling. LUDI addresses these issues by revising the loss function and model structure.

What's New in LUDI
- Implements a less uniform loss to guide each reverse transition more directly towards reconstructing the original (clean) token.
- Adds per-token time embeddings, providing token-level corruption information to the model. This enables confidence-based, fewer-step sampling during generation.
- Demonstrated by continuing training of a 7B autoregressive model to create LUDI-7B, which is capable of complex reasoning and achieves substantial speed improvements.
Performance Improvements
- LUDI-7B achieves a 3-token-per-step decoding speedup compared to standard autoregressive (AR) decoding.
- Maintains competitive performance relative to masked diffusion baselines.
Limitations and Outlook
- Scaling UDLMs remains challenging, since the full potential for complex generation is not yet realized.
