ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research 事件
PRODUCT_LAUNCH2026-06-09影响: MEDIUM
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research arXiv:2606.07591v1 Announce Type: cross Abstract: AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating autonomous scientific research across 40 tasks from 10 scientific domains. Each task is grounded in a real published paper, provides related literature and raw data, and hide
相关产品查看全部 (10)
相关报道查看全部 (1)
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
ArXiv CS.AI2026-06-09