HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection 文章

ArXiv CS.CV2026-08-05PAPERen作者: Jiashu Yao, Haoyu Wen, Siyuan Gao, Yuhang Guo, Zeming Liu, Heyan Huang

详细信息

来源站点
ArXiv CS.CV
作者
Jiashu Yao, Haoyu Wen, Siyuan Gao, Yuhang Guo, Zeming Liu, Heyan Huang
文章类型
PAPER
语言
en
发布日期
2026-08-05

摘要

arXiv:2509.23690v2 Announce Type: replace Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a home and reports hazards before they cause harm. We introduce HomeSafeBench, the first benchmark for free-exploration home safety inspection with egocentric visual feedback, in which an embodied agent navigates a fully interactive 3D home, adjusts its viewpoint, and reports hazards purely from rendered first-person views. Built on the VirtualHome simulator, it covers five categories of common household hazards and comprises 1,000 human-validated inspection tasks. Evaluating a broad range of state-of-the-art Vision-Language Models (VLMs) reveals a large gap, where the best model reaches only about 34.7% F1, far below the 98.0% of a human inspector.