Authors:Xiaozhong Lyu*, Gen Li*, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang (*equal contribution; order interchangeable)
Abstract
ReViV: A feed-forward, unified egocentric 4D reconstruction model, pretrained on more than 600 hours of egocentric video data, that extracts viewer and view dynamics from a single monocular RGB video. It models the joint probability distribution of multimodal signals, including RGB video, camera trajectory, gaze direction, body motion, hand motion, and depth. ReViV achieves SOTA/competitive accuracy and fast inference across various tasks, eliminating the need for task-specific priors from external models.