永不回望:理解自我中心视频中 3D 物体记忆的持久性
Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos
Shravan Chaudhari · William Paul · Suchi Saria · Rama Chellappa · Homanga Bharadhwaj
中文摘要
在日常生活中,我们会遇到一些当时看似无关、之后却变得重要的物体。即使事前并无需求,我们仍能记起某物放在何处或容器内装着什么。本文研究具身助手如何通过观察人的日常活动,从自我中心(egocentric)视频中构建类似的记忆系统,并提出 Ledger——一种持久化的 3D 物体记忆,融合了物体位置、其历史轨迹以及上下文描述。该系统在整个录制过程中关联同一物体的多次观测,并在物体离开视野后仍予以保留,即便人物从未触碰过它。它将每个物体的观测按其静止位置进行聚类,仅在多次重复证据后才记录一次移动,以降低定位噪声的影响。简短的描述则保留诸如容器内物品或支撑面等细节。这些记录可在后续回答空间类问题时直接使用,无需再访问原始图像或视频。在 HD-EPIC 上,该记忆将准确率由 29.7% 提升至 42.6%;在 UCS-Bench 上由 33.8% 提升至 38.5%;在 Ego4D 物体定位任务上,返回预测的中位误差为 0.99 米。分析表明,时间持久性、上下文描述与检索三者起到互补作用。对 100 段由多个场景拼接而成的视频流的进一步实验揭示了检索与构建两个环节的失败模式;逐场景构建可部分恢复因场景切换而损失的性能,缩小其与单场景流之间的差距。
关键要点
- 01问题:具身助手需从日常自我中心视频中构建可回答未来空间查询的长期物体记忆
- 02方法:Ledger 持久化 3D 物体记忆,关联多次观测并按静止位置聚类,多次证据确认后才记录移动,辅以物体及容器内容的简短描述
- 03结果:HD-EPIC 准确率由 29.7% 提升至 42.6%,UCS-Bench 由 33.8% 提升至 38.5%,Ego4D 定位中位误差 0.99 米
- 04分析:时间持久性、上下文描述与检索三者互补,但在多场景拼接视频中存在检索与构建的失败
- 05局限:跨场景切换会带来性能下降,逐场景重建只能部分弥补,完整长程记忆仍存在挑战
解读
尚无解读。
原始英文摘要
arXiv:2610.10538v1 Announce Type: cross Abstract: As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.