vLLM hits 500K GPUs as co-founder Simon Mo makes the case for open models

Share

vLLM, the inference engine born at UC Berkeley’s Sky Computing Lab, now runs on 500,000 GPUs worldwide, according to the project’s own maintainers. The core innovation, called PagedAttention, enables more efficient GPU memory management by optimizing the allocation of the KV cache used by large language models. The project supports over 500 model architectures and more than 200 accelerator types, and became a PyTorch Foundation project in May 2025. Simon Mo, one of the core maintainers, co-founded Inferact in 2025, a company that raised a $150M seed round at an $800M valuation, led by a16z and Lightspeed. Mo argues for open-weight models in production based on three pillars: control, customization, and cost.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles