From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping 文章

ArXiv CS.AI2026-07-02PAPERen作者: Jian Song, Tian Zi, Shen Guanting

详细信息

来源站点
ArXiv CS.AI
作者
Jian Song, Tian Zi, Shen Guanting
文章类型
PAPER
语言
en
发布日期
2026-07-02

摘要

arXiv:2607.00530v1 Announce Type: cross Abstract: Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 15 percentage point gain in end-to-end task success (from 75% in a multimodal baseline system to 90% in an improved configuration identified through a prior ablation study) is sufficient to produce consistent and measurable differences in user perception. The baseline system combines Whisper for speech recognition, Florence-2 for open-vocabulary object detection, LLaMA 3.1 for action extraction, and an interval Type-2 fuzzy logic controller for motion execution. The improved configuration replaces the perception and language modules with Grounding DINO + SAM and Qwen 3.5 9B, respectively, while retaining the same controller.