Google Cloud AI Research, in collaboration with researchers at Washington University in St. Louis and UNC Chapel Hill, launched EnvHarness, a programmable wrapper that transforms static training environments into adaptive ones targeting specific agent weaknesses. On the ALFWorld benchmark, performance rose from 62.4% to 68.3%, a 5.9-point improvement, while out-of-distribution tasks showed a 9.0-point gain reaching 70.4%. On SWE-bench Verified, EnvHarness scored 54.79 compared to 52.13 for the original static environment and 50.37 for environments generated from scratch. Agents trained with this tool used approximately 9.8% fewer interaction steps, reducing compute costs and accelerating training cycles. The tool was open-sourced on GitHub under google-research/envharness.
Source: Read the original article

