InnoEval and new benchmarks show AI models struggle with original research

Share

A series of studies from 2025 and 2026, conducted with the InnoEval evaluation framework, conclude that frontier AI models are significantly more effective at reassembling existing knowledge than inventing new methods. The Reconstruction benchmark, tested across 643 papers, showed that AI models could only recover the central idea 3 to 15 percent of the time. A study published in Science in October 2026 compared AI-generated hypotheses to human ones, with AI scoring an average novelty of 3.406 compared to 3.968 for humans. Another large-scale study of 121,000 preprints found that LLMs tend to generate narrowly focused ideas with little diversity. These results suggest that current AI models are better suited as research assistants than as sources of original hypotheses.

Source: Read the original article

Disclaimer: this content is for information purposes only and is not financial advice. Cryptocurrencies are highly volatile: you may lose all of your capital. Always do your own research. Legal notice
Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Read More

Items