详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Andy J. Phu, James Mooney, Karin de Langis, Khanh Chi Le, Dongyeop Kang
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-08-03
摘要
arXiv:2607.29181v1 Announce Type: cross Abstract: Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building these models from raw, unstructured screen activity remains an open challenge. We present SERUM, a multi-pass framework that extracts finite-state behavioral models directly from unstructured egocentric video using hierarchical VLM annotation. Processing screen recordings through a sliding window, SERUM alternates between activity-recognition and intent-inference passes, with each pass refining labels using accumulated prior context to reduce hallucination and temporal conflation seen in single-pass annotation. Synonymous states are then merged via sentence embeddings and human-calibrated thresholds into a compact, coherent taxonomy.