ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research 事件

PRODUCT_LAUNCH2026-06-09影响: MEDIUM

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research arXiv:2606.07591v1 Announce Type: cross Abstract: AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating autonomous scientific research across 40 tasks from 10 scientific domains. Each task is grounded in a real published paper, provides related literature and raw data, and hide