RayTun3R: Online Camera Adaptation in 3D Foundation Models 文章

ArXiv CS.CV2026-07-07PAPERen作者: Daniil Sinitsyn, Nikita Araslanov, Daniel Cremers

详细信息

来源站点
ArXiv CS.CV
作者
Daniil Sinitsyn, Nikita Araslanov, Daniel Cremers
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.02711v1 Announce Type: new Abstract: Recent 3D foundation models, such as DUSt3R, MASt3R, VGGT, $\pi^3$, and Depth Anything 3, provide strong feed-forward depth and pose estimates on pinhole imagery, but degrade sharply under fisheye camera geometry. We show that this failure is partly caused by a pinhole camera bias in the positional encodings of pretrained 3D foundation models, and propose RayTun3R, a lightweight camera adaptation approach. It keeps the pretrained network fixed and adapts only lightweight components tied to token position and camera geometry. RayTun3R learns parameter-efficient residual corrections to absolute and rotary positional encodings, together with parameter-free tokenization and corrections to prediction-grid coordinates that remove residual pinhole assumptions. The resulting adapter contains only 10,752 trainable parameters and can be learned from a short temporal segment using geometric losses.