Microsoft researchers published a paper on September 10 introducing a technique called environment-probing curation, which allows AI agents to verify their own memories against the real environment before storing them. On the CLBench database exploration benchmark, the pass rate jumped from 39% to 73%, while the cost per question was halved from $3.38 to $1.68. The number of queries per question dropped from 8.8 to 4.7, and tool calls declined between 16% and 75% depending on the task. Tested on six APEX management consulting scenarios with multiple foundation models including Sonnet 4.6 and Opus 4.7, this system requires no schema changes or retraining of the underlying agent. The testing environment used was GitHub Copilot, one of Microsoft’s flagship AI products already deployed across millions of developer workflows.
Source: Read the original article

