ECHO:人体物体携带行为的具身相机观测数据集
ECHO: Embodied Camera Observations of Human Object Carrying
Xuefei Sun · Lorin Achey · Kali Hamilton · Alberto Speranzon · Gregory Grebe · Yonatan Bisk · et al.
中文摘要
具身智能体与辅助智能体不仅需要识别物体,还需根据环境布局与居住者习惯推理物体的合理放置位置。现有 RGB-D 扫描数据集重建静态房间但缺乏人体活动记录,而人-物交互数据集虽然捕捉运动却缺少可导航的完整重建场景及物体自然目的地的真值标注。为解决这一问题,论文提出上下文物体放置(contextual object placement)基准任务:在观测到的物体携带事件中预测物体的目的地位置,并构建了 ECHO(Embodied Camera observations of Human Object carrying)大规模合成数据集。ECHO 将室内场景的稠密 RGB-D 扫描与具身人体携带日常物品前往符合上下文的目的地的录像配对,是首个同时包含重建场景、人体活动、自然语言与上下文放置标注的公开数据集,涵盖 115 个 HM3D 场景的 159 层楼面,共 3,805 条人工标注事件,涉及 198 个不同物体。每层楼面提供完整 RGB-D 扫描及人工标注的房间标签与表面列表;每条事件包含同步的 RGB-D 遭遇片段、6-DoF 相机/人体/物体轨迹、起止表面、动作描述及一条人工撰写的上下文语句。基于输入掩蔽探针与端到端基线的实验表明,没有任何单一输入模态足够完成该任务,凸显了对场景结构、人体活动与上下文知识进行联合推理的必要性。
关键要点
- 01问题:具身/辅助智能体缺乏对物体在环境中合理放置位置的推理能力,且缺少兼具重建场景、人体活动与自然语言标注的公开基准数据集
- 02方法:提出上下文物体放置任务,构建 ECHO 合成数据集,配对稠密 RGB-D 室内扫描与具身人体携带物体至上下文合理目的地的记录
- 03数据规模:3,805 条人工标注事件,跨 115 个 HM3D 场景的 159 层楼面,涉及 198 个物体,每事件含 RGB-D 片段、6-DoF 轨迹、上下文语句等
- 04结果:输入掩蔽探针与端到端基线评估显示无单一模态足以解决该任务,需联合利用场景结构、人体活动与上下文知识
- 05局限/开放方向:联合多模态推理仍是未解决问题,需探索融合场景、人体动作与语言上下文的模型
解读
尚无解读。
原始英文摘要
arXiv:2610.10438v1 Announce Type: cross Abstract: Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it. Progress on this problem has been limited, in part because no dedicated benchmark or dataset exists to define and evaluate it. Existing RGB-D scan datasets reconstruct static rooms without human activity, while human-object-interaction datasets capture motion without a navigable, fully reconstructed scene or a ground-truth notion of an object's natural destination. We introduce contextual object placement as a benchmark task: predicting an object's destination during an observed object-carrying episode. To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations. ECHO is the first publicly available dataset to combine reconstructed scenes, human activity, natural language, and contextual-placement annotations. It comprises 3,805 human-annotated episodes across 159 floors of 115 HM3D scenes, involving 198 distinct objects. Each floor includes a complete RGB-D scan with human-annotated room labels and a surface list. Each episode provides synchronized RGB-D encounter clips; 6-DoF camera, human, and object trajectories; start and destination surfaces; an action caption; and a human-written context: a single sentence describing the inhabitant's routine that implies the destination without naming it. We evaluate contextual object placement using input-masked probes and an end-to-end baseline. Results show that no single input modality is sufficient, highlighting the need to jointly reason over scene structure, human activity, and contextual knowledge.