论文
arXiv
SpatialIntelligence
Trajectory
Mobility
LLM
Multimodal
UrbanTraffic
GeoSimulation
中文标题
SIREN-Bench:面向行为的紧急车辆交互生成与评估
English Title
SIREN-Bench: Behavior-Driven Generation and Evaluation of Emergency-Vehicle Interactions
Yicheng Zhu, Tianmu Zhao, Haoxin Leng, Fan Zuo, Tao Li, Zilin Bian
发布时间
2026/8/25 13:41:34
来源类型
preprint
语言
en
摘要
中文对照

紧急车辆(EMV)可通过促使民用车辆制动、变道或形成救援通道等方式,重构周边交通流。评估此类安全关键型交互需在行为层面同时控制EMV的通行特权与民用车辆的响应,并具备一致的感知能力与真值标注。现有数据集与仿真基准未直接提供该组合能力。本文提出\textbf{SIREN}——一种面向行为的SUMO–CARLA协同仿真平台,用于生成EMV–民用车辆交互。SIREN将SUMO的路网级交通演化与行为逻辑,与CARLA的连续车辆控制及同步车载感知相结合;交互控制依据激活的行为模式,由SUMO、CARLA或二者协同执行。我们基于该平台构建\textbf{SIREN-Bench-v1},涵盖七个参数化交互模板,覆盖紧急等级L1–L3及三类行为族,并提供同步传感器观测与仿真器原生标注。我们通过三项典型任务展示该基准:3D目标检测、轨迹预测与视觉-语言风险理解。对九种轨迹预测器、四种基于LiDAR的检测器及五种视觉-语言模型的评估揭示了行为依赖的失效模式:交通清空类交互对检测最具挑战性,享有通行特权的路口穿越对预测最具挑战性,且无学习型预测器在平均意义上优于恒速参考模型;视觉-语言模型在正常交通场景下的表现显著优于近碰撞与碰撞事件场景。上述结果验证了以行为为中心的基准测试价值,并确立SIREN作为面向自动驾驶的可扩展数据生成与评估平台。

English Original

Emergency vehicles (EMVs) can reorganize surrounding traffic as civilian vehicles brake, change lanes, or form rescue corridors in response to their passage. Evaluating these safety-critical interactions requires behavior-level control over both EMV privileges and civilian responses, together with consistent sensing and ground truth. Existing datasets and simulation benchmarks do not directly provide this combination. We present \textbf{SIREN}, a behavior-driven SUMO--CARLA co-simulation platform for generating EMV--civilian interactions. SIREN couples SUMO's network-level traffic evolution and behavior logic with CARLA's continuous vehicle control and synchronized onboard sensing; depending on the active behavior, the interaction is controlled by SUMO, CARLA, or jointly. We instantiate the platform as \textbf{SIREN-Bench-v1}, comprising seven parameterized interaction templates across emergency levels L1--L3 and three behavior families, with synchronized sensor observations and simulator-native annotations. We demonstrate the benchmark through three representative tasks: 3D object detection, trajectory prediction, and vision-language risk understanding. Evaluations of nine trajectory predictors, four LiDAR-based detectors, and five vision-language models reveal behavior-dependent failure modes. Traffic-clearance interactions are hardest for detection, privileged intersection traversal is hardest for prediction, and no learned predictor outperforms the constant-velocity reference on average. Vision-language models perform substantially better on normal traffic than on near-miss and collision events. These results demonstrate the value of behavior-centered benchmarking and establish SIREN as an extensible data-generation and evaluation platform for autonomous-driving and transportation safety research.

我的阅读记录

正在加载阅读记录…

元数据
arXiv2608.24094v1
来源arXiv
类型论文
抽取状态raw
关键词
SpatialIntelligence
Trajectory
Mobility
LLM
Multimodal
UrbanTraffic
GeoSimulation
cs.RO