The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence 文章

ArXiv CS.CL2026-08-03PAPERen作者: David Nordfors

详细信息

来源站点
ArXiv CS.CL
作者
David Nordfors
文章类型
PAPER
语言
en
发布日期
2026-08-03

摘要

arXiv:2606.21008v2 Announce Type: replace Abstract: The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs. No content is given in advance; the contestants create all of it -- a new kind of analogy test, analogical production falsifiable sentence by sentence, with no fixed test set to leak into training (contamination-resistant by construction). In the council-of-peers benchmark, the contestants also rate each other's creations. We introduce the first spectral solution, to our knowledge, to the wicked problem of benchmarking LLMs' factual accuracy without golden keys or oracle models: one singular value decomposition of the evaluators' ratings matrix yields their competence as both generators and judges of true statements at once. Competence on the subjective criteria comes from each judge's rating consistency as the yardstick shifts. The factual rating correlates with GPQA Diamond at Pearson r = 0.

相关事件

暂无数据

相关公司查看全部 (6)

A
ACLUNIVERSITY
A
AMI团队RESEARCH_INSTITUTE
A
ActuaNONPROFIT

相关人物

暂无数据