PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models 事件

Name: PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models
Start: 2026-06-03

PRODUCT_LAUNCH2026-06-03影响: MEDIUM

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models arXiv:2606.03858v1 Announce Type: new Abstract: Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and mathematical reasoning, hindering the interpretability of failures in math tasks. We introduce PyraMathBench, a comprehensive hierarchical bench

人工智能

关系图谱

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models 事件

相关公司查看全部 (10)

相关人物查看全部 (1)

相关产品查看全部 (10)

相关技术查看全部 (10)

相关报道查看全部 (1)