Three US AI infrastructure providers, Modal, Fireworks AI, and Baseten, are now offering hosted inference for Moonshot AI’s Kimi K3 at roughly one-tenth the cost of direct access. This 2.8 trillion parameter model, based on a Mixture-of-Experts architecture, benefits from Nvidia GB300 and AMD MI350X and MI355X accelerators, which are restricted in China due to US export controls. Modal achieves 460 tokens per second with its custom speculative decoder called DFlash. Pricing stands at approximately $3 per million input tokens, $0.30 per million cached tokens, and $15 per million output tokens. The providers emphasize zero-data retention policies and OpenAI-compatible APIs as key enterprise selling points.
Source: Read the original article

