Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty? 文章

ArXiv CS.AI2026-07-08PAPERen作者: Satiyabooshan Murugaboopathy, Connor T. Jerzak, Adel Daoud

详细信息

来源站点
ArXiv CS.AI
作者
Satiyabooshan Murugaboopathy, Connor T. Jerzak, Adel Daoud
文章类型
PAPER
语言
en
发布日期
2026-07-08

摘要

arXiv:2508.01109v3 Announce Type: replace Abstract: We investigate whether socioeconomic indicators, like household wealth, leave recoverable informational imprints in both satellite imagery (capturing features like buildings and roads) and Internet-sourced text (reflecting historical, cultural, and narratives of neighborhoods). Using DHS data from African neighborhoods (clusters), we pair high-resolution Landsat images with textual descriptions generated by LLMs conditioned on location/year, plus text retrieved by an LLM-driven AI Search Agent from web sources. We develop a multimodal framework that predicts household wealth (International Wealth Index; IWI) via five pipelines: (i) a vision model on satellite images, (ii) an LLM using only location and year, (iii) an AI agent that searches and synthesizes web text, (iv) a joint image-text encoder, and (v) an ensemble of all signals. Our framework yields three contributions.