Alibaba launched Qwen3.8-Max on August 3, a massive 2.4 trillion parameter model capable of processing over 4,000 tokens per second per GPU on Nvidia’s GB300 NVL72 hardware, generating approximately 3,000 words every second on a single chip. The model operates as a sparse Mixture-of-Experts multimodal system natively handling text, images, and video with a 1 million token context window. On the OSWorld-Verified benchmark, Qwen3.8-Max scored 86.1, outperforming Anthropic’s Claude Fable 5 (85.0) and OpenAI’s GPT-5.6 Sol Max, ranking as the second highest-performing AI model behind Fable 5. Alibaba is pricing API access at $2 per million input tokens and $6 per million output tokens, and plans to release open weights for both Qwen3.8-Max and its smaller variant Qwen3.8-27B.
Source: Read the original article

