Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment 文章

ArXiv CS.CV2026-08-04PAPERen作者: Vishwajeet Shivaji Hogale, Anjali Pai, Nitya Ravi

详细信息

来源站点
ArXiv CS.CV
作者
Vishwajeet Shivaji Hogale, Anjali Pai, Nitya Ravi
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.02470v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remains unreliable for fine-grained, visually ambiguous targets. We study this gap in the context of automated vehicle damage assessment, where fine-grained defects such as scratches and hairline cracks occupy few pixels, produce weak gradient signal, and are easily confused with reflections and surface texture. We show that a state-of-the-art VLM (Qwen-VL) achieves strong semantic classification accuracy (87.3%) on this task but is systematically ungrounded at the spatial level: it hallucinates damage in reflective regions, misses elongated scratches entirely, and produces spatially inconsistent outputs when prompted for localization.

相关事件

暂无数据

相关公司查看全部 (2)

A
AT TCOMPANY

相关人物

暂无数据