Nvidia published a research paper on August 20 introducing the ACES framework (Agentic Continuous Evaluation of Skills), revealing that the correlation between traditional scan-only metrics and LLM-judge scores is only 0.14 (Spearman rho), representing near-zero predictive power for actual runtime performance. ACES measures « Skill Lift » by comparing agent performance with and without a given skill enabled. Tested across 947 paired cases spanning 58 of 64 production skills, the framework shows a mean Skill Lift of 0.2134 with 72.8% of cases displaying positive lift. Nvidia released SkillEvaluator as an open-source tool, integrating the Tier 3 live evaluation component into its Verified Agent Skills pipeline for quality control of enterprise AI skills.
Source: Read the original article

