Inception, an AI startup, has launched Mercury 2.5, a diffusion-based large language model offering a 40% intelligence improvement over its predecessor while processing 1,107 tokens per second on standard NVIDIA GPUs. The model features a 260,000-token context window with standard pricing of $0.20 per million input tokens and $0.75 per million output tokens. The company raised $50 million in funding led by Menlo Ventures with backing from Andrew Ng and Andrej Karpathy. The model ships via an OpenAI-compatible API with an 80% launch discount bringing costs down to $0.04 per million input tokens. Early adopter OpenCall reported median latencies under 200 milliseconds for voice interactions.
Source: Read the original article

