在线多相机 3D 跟踪需在同步视图中维持场景全局身份,但基于查询的跟踪器仅在实例库中隐式携带这些身份,导致查询中断时身份碎片化。本文提出一种在线架构,通过在循环稀疏查询上显式预测 ID 来恢复关联精度。采用由外向内的 Sparse4D 检测器将校准视图融合为世界坐标系下的 3D 检测结果,并传播稀疏查询库;同时利用因果 MOTIP ID 解码器在有限轨迹记忆中对检测结果进行关联。我们将 MOTIP 的相对 ID 预测和回收槽位运行时机制适配至全局融合的 3D 观测数据,并引入度量空间门控及基于邻近度的新生目标恢复机制。在官方 2026 AI City Challenge Track 1 测试集上,该方法将 HOTA 从使用原生实例库身份的 29.63 提升至 38.01,主要得益于 AssA 从 20.83 增至 31.10,并在公开排行榜中位列第三。对每个场景全部 9,000 帧进行的完整序列验证表明,解耦 ID 训练相较于原生身份能提升 HOTA,而在分离 ID 目标的同时继续训练检测器则会产生依赖于场景的增益与损失。
Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only implicitly in the instance bank, where they fragment upon query interruption. We present an online architecture that recovers association accuracy by predicting IDs explicitly over recurrent sparse queries. An outside-in Sparse4D detector fuses calibrated views into world frame 3D detections while propagating a sparse query bank, and a causal MOTIP ID decoder associates detections against a finite trajectory memory. We adapt MOTIP's relative-ID prediction and recycled slot runtime to globally fused 3D observations, and introduce metric spatial gating and proximity based newborn recovery. On the official 2026 AI City Challenge Track 1 test set, our method raises HOTA from 29.63 with native instance bank identities to 38.01, primarily through an AssA increase from 20.83 to 31.10, and ranks third on the public leaderboard. Full-sequence validation over all 9,000 frames of each scene shows that decoupled ID training improves HOTA over native identities, whereas continuing detector training alongside the detached ID objective produces scene-dependent gains and losses.