跳到正文
返回论文列表
cs.CV提交于 已译

MOTIP2:面向端到端多目标跟踪的空间先验

MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking

Beno\^it Roussel · Damien Bouet · Liming Chen · Pierre Perrault

中文摘要

端到端多目标跟踪器已在关联困难基准上缩小了与经典跟踪-检测范式的差距,但仍会犯一些经典跟踪器绝不会犯的、在空间上不合理的错误,例如将同一身份分配给画面两侧的物体。模型本可学会规避这些错误,但跟踪标注稀缺,因此本文显式编码空间先验,同时保持完全端到端推理、无事后关联。提出了三类空间先验,分别作用于数据、损失与表示阶段。空间 ID 切换(Spatial ID Switches)在轨迹置换时偏向空间上重叠的物体,缩小训练与推理混淆之间的错配;空间 ID 损失(Spatial ID Loss)按框距离缩放每个身份的惩罚,使远距离切换代价更高;空间锚点(Spatial Anchor)为每个轨迹 token 注入帧位置,作为显式的空间线索。三类先验在 MOTIP2 中实例化,该方法由 MOTIP 适配并构建于实时 DEIM 检测 Transformer 之上。在不使用额外数据训练时,主模型 MOTIP2-L 创下新的最优水平:DanceTrack 上 73.4 HOTA,SportsMOT 上 76.0,PersonPath22 上 71.1 IDF1。MOTIP2 是一个涵盖速度-精度权衡的模型族:轻量版 MOTIP2-S 以超过 3 倍速度匹配原 MOTIP,MOTIP2-X 在 DanceTrack 上达到 74.8 HOTA。

关键要点

  1. 01问题:端到端多目标跟踪器仍会出现空间不合理错误(如跨画面两侧同 ID),且跟踪标注稀缺。
  2. 02方法:在数据、损失、表示三阶段显式编码空间先验,即 Spatial ID Switches、Spatial ID Loss 与 Spatial Anchor。
  3. 03结果:MOTIP2-L 不使用额外数据,在 DanceTrack、SportsMOT、PersonPath22 创 SOTA;MOTIP2-S 以 3 倍速度匹配原 MOTIP;MOTIP2-X 在 DanceTrack 达 74.8 HOTA。
  4. 04局限:依赖显式空间先验弥补标注不足,未引入额外训练数据。

解读

尚无解读。

原始英文摘要

arXiv:2610.10391v1 Announce Type: new Abstract: End-to-end multi-object trackers have narrowed the gap with classical tracking-by-detection on association-difficult benchmarks. Yet they still make spatially implausible errors no classical tracker would, such as assigning one identity to objects on opposite sides of the frame. A model could learn to avoid them, but tracking annotations are scarce, so we encode spatial priors explicitly instead, while keeping inference fully end-to-end with no post-hoc association. We propose three spatial priors, at the data, loss, and representation stages. Spatial ID Switches bias trajectory permutations toward spatially overlapping objects, reducing the mismatch between training and inference confusions. Spatial ID Loss scales each identity's penalty by its box distance, so a distant switch costs more than a nearby one. Spatial Anchor gives each track token its frame position, an explicit spatial cue for attention. We instantiate the three priors in MOTIP2, a tracker adapted from MOTIP and built on the real-time DEIM detection transformer. Trained without extra data, its main model, MOTIP2-L, sets a new state of the art: 73.4 HOTA on DanceTrack, 76.0 on SportsMOT, and 71.1 IDF1 on PersonPath22. MOTIP2 is a family of models spanning the speed-accuracy trade-off: a lighter model, MOTIP2-S, matches the original MOTIP at over 3x the speed, and MOTIP2-X reaches 74.8 HOTA on DanceTrack.

同方向论文 · cs.CV

查看全部 →