K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos 文章

ArXiv CS.CV2026-07-07PAPERen作者: Khush Attarde, Yusuf Ali, Megha Thukral, Divye Bhutani, Thomas Ploetz, Zsolt Kira

详细信息

来源站点
ArXiv CS.CV
作者
Khush Attarde, Yusuf Ali, Megha Thukral, Divye Bhutani, Thomas Ploetz, Zsolt Kira
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.02680v1 Announce Type: new Abstract: MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies in understanding and modeling animal-centric scenarios. As animals are integral to millions of households, benchmarking next-generation AI models on pet-focused tasks, ranging from recognizing distress signals to enabling responsive robotic companions, is essential for building AI systems that can work alongside us. We introduce K9-Bench, a novel benchmark focused on real-world domestic dog videos, specifically targeting canine action and interaction understanding via approximately 5000 question-answer pairs across 907 videos spanning 5 distinct task categories that test long-form, canine-centric multimodal reasoning in MLLMs.

相关事件

暂无数据

相关公司查看全部 (3)

A
ACTIONNONPROFIT
A
ANDINONPROFIT
A
AT TCOMPANY

相关人物

暂无数据