An Intelligent-Cloud Edge Multimodal Interaction System for Robots 文章

ArXiv CS.AI2026-07-17PAPERen作者: Zihan Guo, Xiaoqi Li

详细信息

来源站点
ArXiv CS.AI
作者
Zihan Guo, Xiaoqi Li
文章类型
PAPER
语言
en
发布日期
2026-07-17

摘要

arXiv:2607.14675v1 Announce Type: cross Abstract: Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据