LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability 文章

ArXiv CS.AI2026-07-03PAPERen作者: Aastha Sapkota, M. G. Sarwar Murshed

详细信息

来源站点
ArXiv CS.AI
作者
Aastha Sapkota, M. G. Sarwar Murshed
文章类型
PAPER
语言
en
发布日期
2026-07-03

摘要

arXiv:2607.01247v1 Announce Type: cross Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate steps. They are also difficult to grade at scale because instructors must apply partial-credit rubrics consistently while giving feedback that helps students repair misconceptions. This paper evaluates six contemporary large language model (LLM) configurations, Gemini 3.1 Pro Extended, Gemini 3.5 Flash, ChatGPT 5.5 Pro Extended, ChatGPT 5.5 Thinking, Claude Pro Opus 4.7, and Claude Sonnet 4.6, as grading assistants for an undergraduate discrete mathematics examination. The study compares two grading policies. The BASELINE policy uses a stricter rubric-following prompt that emphasizes explicit evidence and complete justification.