FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling 文章

ArXiv CS.CV2026-08-03PAPERen作者: Li Lin, Wujun Xu, Weiwei Meng, Kaiwen Xia, Kang Hao Cheong, Shuai Wang

详细信息

来源站点
ArXiv CS.CV
作者
Li Lin, Wujun Xu, Weiwei Meng, Kaiwen Xia, Kang Hao Cheong, Shuai Wang
文章类型
PAPER
语言
en
发布日期
2026-08-03

摘要

arXiv:2607.29596v1 Announce Type: cross Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes an efficiency-decreasing issue. To reconcile the conflict between capturing temporal information and maintaining inference efficiency in VLAs, this paper introduces FibVLA, an efficient framework featuring temporal perception of long-context history. Specifically, we leverage logarithmic hindsight sampling to both proprioceptive states and visual frames to capture long-term temporal dependencies with minimal redundancy.

相关事件

暂无数据

相关人物

暂无数据