A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions 文章

ArXiv CS.CV2026-07-16PAPERen作者: Ji Zhou, Yilin Ding, Yongqi Zhao, Jiachen Xu, Dong Bi, Johannes Betz, Arno Eichberger

详细信息

来源站点
ArXiv CS.CV
作者
Ji Zhou, Yilin Ding, Yongqi Zhao, Jiachen Xu, Dong Bi, Johannes Betz, Arno Eichberger
文章类型
PAPER
语言
en
发布日期
2026-07-16

摘要

arXiv:2601.22830v2 Announce Type: replace Abstract: Reliable environmental perception remains one of the main obstacles for safe operation of automated vehicles. Safety of the Intended Functionality (SOTIF) concerns safety risks from perception insufficiencies, particularly under adverse conditions where conventional detectors often falter. While Large Vision-Language Models (LVLMs) demonstrate promising semantic reasoning, their quantitative effectiveness for safety-critical 2D object detection is underexplored. This paper presents a systematic evaluation of ten representative LVLMs using the PeSOTIF dataset, a benchmark specifically curated for long-tail traffic scenarios and environmental degradations. Performance is quantitatively compared against two specialized detectors: the anchor-based YOLOv5 and the transformer-based RT-DETRv4. Experimental results reveal a critical trade-off: top-performing LVLMs (e.g.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据