A study by Stanford’s Hazy Research group reveals that local AI efficiency, measured in intelligence per joule, improved 18-fold in roughly 16 months, from mid-2024 to late 2025. This progress stems from a combination of factors: model architecture improvements contributed roughly 3.1x of the gain, while hardware advances delivered approximately 5.9x. Tests on various chips, including the NVIDIA B200 and SambaNova SN40L, showed that purpose-built AI accelerators significantly outperform consumer-grade silicon on efficiency metrics. Adopting a hybrid local-cloud approach could reduce energy consumption and costs by 60-80% compared to fully cloud-based infrastructure. Local AI systems now achieve approximately 88.7% accuracy on single-turn chat and reasoning queries.
Source: Read the original article

