Coinbase reported on October 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test conducted for its Onramp service, despite an unchanged decision policy. The evaluation covered 16,140 transactions involving 7,293 users, including 813 confirmed fraudulent transactions. Sonnet’s detection rate fell by 22.2 percentage points while its dollar-weighted detection dropped by 22.9 points. GPT’s precision increased by 11.5 points, but its detection rate declined by 20.7 points, illustrating how an isolated improvement in one metric can mask weaker overall performance. A customized Qwen3.5-9B model subsequently outperformed Opus 4.5, with a 9.6 percentage point improvement in its F1 score and a 55% reduction in latency.
Source: Read the original article

