Uno achieves 2.5x higher throughput in LLMs by bolting diffusion onto existing models

Share

Uno is a method that adds lightweight diffusion adapters to existing autoregressive large language models, enabling them to generate multiple candidate tokens in parallel without sacrificing output quality. The Uno-Qwen 8B model, built on top of Qwen’s 8 billion parameter model, achieves approximately 2.5x speedup at batch size 1, with a potential ceiling of up to 3x compared to the base autoregressive model. Developed by Subham Sekhar Sahoo and collaborators, the paper titled « Unlocking Lossless Speedups in LLMs via Discrete Diffusion » was published on September 3, 2026. Code and model weights are available as open source on GitHub and Hugging Face, but these results have not yet undergone peer review. If validated in production, this technique could serve the same volume of requests with roughly 40% fewer GPUs, without requiring organizations to replace their existing autoregressive infrastructure.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles