Alibaba’s Qwen3.8 model goes live on Nvidia’s GB300, hits 4,000 tokens per second

Share

Alibaba launched Qwen3.8-Max on August 3, a massive 2.4 trillion parameter model capable of processing over 4,000 tokens per second per GPU on Nvidia’s GB300 NVL72 hardware, generating approximately 3,000 words every second on a single chip. The model operates as a sparse Mixture-of-Experts multimodal system natively handling text, images, and video with a 1 million token context window. On the OSWorld-Verified benchmark, Qwen3.8-Max scored 86.1, outperforming Anthropic’s Claude Fable 5 (85.0) and OpenAI’s GPT-5.6 Sol Max, ranking as the second highest-performing AI model behind Fable 5. Alibaba is pricing API access at $2 per million input tokens and $6 per million output tokens, and plans to release open weights for both Qwen3.8-Max and its smaller variant Qwen3.8-27B.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles