Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models 文章

ArXiv CS.CV2026-07-28PAPERen作者: Huafu Li, Guo Chen, Jia Xia, Lei Wang, Wei Du, Yun Yao, Weijun Peng, Liming Li

详细信息

来源站点
ArXiv CS.CV
作者
Huafu Li, Guo Chen, Jia Xia, Lei Wang, Wei Du, Yun Yao, Weijun Peng, Liming Li
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.22723v1 Announce Type: new Abstract: Visual information extraction (VIE) from visually rich documents remains challenging due to high layout variability and real-world impairments. Existing methods typically rely on sequential OCR pipelines or end-to-end models requiring extensive labeled data and layout-specific training, limiting their scalability.We propose a classification-guided large vision-language model (LVLM) framework for multi-type VIE that achieves high accuracy with minimal supervision. The approach decouples document-type classification from content extraction and employs in-context learning (ICL)-based dynamic prompt engineering to inject task-specific knowledge, enabling robust zero-shot inference across diverse layouts. From a theoretical perspective, the proposed method can be viewed as a form of conditional computation that reduces task uncertainty and improves information efficiency during prompt-based inference.

相关事件

暂无数据

相关公司查看全部 (3)

A
AMI团队RESEARCH_INSTITUTE
A
ACTIONNONPROFIT

相关人物

暂无数据