vLLM, the inference engine born at UC Berkeley’s Sky Computing Lab, now runs on 500,000 GPUs worldwide, according to the project’s own maintainers. The core innovation, called PagedAttention, enables more efficient GPU memory management by optimizing the allocation of the KV cache used by large language models. The project supports over 500 model architectures and more than 200 accelerator types, and became a PyTorch Foundation project in May 2025. Simon Mo, one of the core maintainers, co-founded Inferact in 2025, a company that raised a $150M seed round at an $800M valuation, led by a16z and Lightspeed. Mo argues for open-weight models in production based on three pillars: control, customization, and cost.
Source: Read the original article

