论文
arXiv
GeoAI
GIS
SpatialIntelligence
Trajectory
Mobility
LLM
Multimodal
GeoMultimodal
UrbanTraffic
中文标题
VANTAGE-Bench:评估视觉-语言模型中的基础设施人工智能差距
English Title
VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
发布时间
2026/9/9 03:50:25
来源类型
preprint
语言
en
摘要
中文对照

随着视觉-语言模型(VLM)逐步迈向物理场景部署,当前研究重心仍集中于面向动作的具身人工智能(Embodied AI),其评测主要基于以主体为中心的消费级视频。这一取向忽视了一类广泛存在的物理人工智能——基础设施人工智能(Infrastructure AI),该范式依赖固定摄像头,以开环方式提供安全监控、运行日志等洞察。为此,我们提出VANTAGE-Bench基准,用于量化这一“基础设施人工智能差距”。该基准覆盖物流、交通与智能空间三大运营领域,统一涵盖图像与视频在语义、空间、时间及时空联合维度的能力评估,并突破多项选择题范式,引入包括密集视频描述与时空定位在内的八种任务形式;针对单目标跟踪(Single Object Tracking),新增单次前向轨迹协议(single-pass trajectory protocol),并据我们所知,首次在固定摄像头基础设施视频上开展此类评测,且以专业追踪器为评分基准。标注覆盖三类范式,总计3,346个媒体资产:3,342条视频-任务标注、4,281条图像-定位标注,以及27,404个检测框。对17个模型进行零样本评测后发现,相较于以消费者为中心的基准,性能缺口呈现集中性而非普遍性:事件验证、指代表达与时间定位任务在所有模型规模下均落后约9至24个百分点,而视频问答任务与VideoMME的差距则控制在5.3分以内,2D空间指向任务亦未表现出相对于BLINK的性能缺口。时间维度能力在绝对指标上最弱:无一系统在时间定位任务上的mIoU超过55.7,在密集视频描述任务上的SODA_c得分超过37.3。在跟踪任务中,前沿模型在短时域内与专业追踪器相差约5分,但随预测时域延长,性能差距显著扩大。开源权重模型在2D目标定位任务上全面领先,表明性能缺口既非单纯由模型规模亦非仅由闭源属性所致。

English Original

As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overlooks a pervasive class of Physical AI: Infrastructure AI, which relies on fixed cameras for open-loop insights like safety monitoring and operational logging. We introduce VANTAGE-Bench, a benchmark measuring this "Infrastructure AI Gap." It spans three operational domains (Logistics, Transportation, and Smart Spaces), unifies image and video evaluation across semantic, spatial, temporal, and spatio-temporal capabilities, and moves beyond multiple-choice to eight task formulations including dense captioning and spatio-temporal grounding. It adds a single-pass trajectory protocol for Single Object Tracking and, to our knowledge, the first such evaluation on fixed-camera infrastructure video, scored against specialist trackers. Annotation spans three regimes over 3,346 media assets: 3,342 video-task annotations, 4,281 image-grounding annotations, and 27,404 detection boxes. Evaluating 17 models zero-shot, we find the shortfall relative to consumer-centric benchmarks is concentrated, not general. Event verification, referring expressions, and temporal localization fall roughly 9 to 24 points at every model scale, while video question answering stays within 5.3 points of VideoMME and 2D spatial pointing shows no shortfall against BLINK. The temporal pillar is weakest in absolute terms: no system exceeds 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning. On tracking, frontier models come within roughly 5 points of specialist trackers over short horizons but separate as the horizon extends. Open-weight models lead 2D object localization outright, so neither scale nor proprietary access explains the pattern. Data, evaluation harness, and leaderboard: https://vantage-bench.org/

我的阅读记录

正在加载阅读记录…

元数据
arXiv2609.09396v1
来源arXiv
类型论文
抽取状态raw
关键词
GeoAI
GIS
SpatialIntelligence
Trajectory
Mobility
LLM
Multimodal
GeoMultimodal
UrbanTraffic
cs.CV
cs.AI