Chutes AI, a decentralized AI inference platform running on Bittensor’s Subnet 64, partnered with Harvard University researchers to release what may be the largest public dataset of LLM serving metadata ever assembled. The dataset covers a full year from April 11, 2025 to April 12, 2026, comprising 6.12 billion requests from 314,970 anonymized users across 9,174 models, totaling 35.8 trillion input tokens and 2.52 trillion output tokens. A key finding reveals that 99% of repeat requests occur within a 15-minute window, suggesting that prefix-aware routing strategies could achieve near-optimal cache hit rates. Researchers also note a gradual decline in output length, dropping from hundreds of tokens to fewer than 100, indicating a shift toward automated or machine-driven queries. The dataset is now freely available through GitHub and a Harvard S3 bucket.
Source: Read the original article

