车联万物(V2X)协同可实现超视距感知,缓解单车传感器因遮挡导致的感知局限。然而,现有V2X基准对闭环评估和语言引导监督的支持有限,制约了面向端到端协同驾驶的视觉-语言模型(VLM)的发展。为解决上述局限,我们提出V2XBench——一个具备同步主车-路侧感知与闭环评估能力的仿真平台,以及Chat-V2XBench——一个面向协同推理、渐进式结构化的视觉问答(VQA)数据集。基于该基准基础设施,我们进一步提出AURORA,一种端到端协同驾驶框架。该框架采用双视角感知架构,并通过查询级跨视角查询对齐与融合(CQAF)模块,在查询层面缓解主车与路侧视角间的空间与语义差异。利用由此生成的统一表征令牌,一个经LoRA适配的VLM实现了语义推理与生成式轨迹规划的联合建模。在V2XBench上开展的大量闭环评估表明,AURORA在严重遮挡场景下达到当前最优性能,路径完成率达98.21%,驾驶评分为76.02,同时仅需较低的路侧通信带宽。本工作最终开创了一种可扩展的V2X-VLM范式,为下一代协同自动驾驶奠定基础。
Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and language-grounded supervision, hindering the development of vision-language models (VLMs) for end-to-end cooperative driving. To address these limitations, we introduce V2XBench, a simulation platform featuring synchronized ego--roadside sensing and closed-loop evaluation, together with Chat-V2XBench, a progressively structured VQA dataset for cooperative reasoning. Building upon this benchmark infrastructure, we propose AURORA, an end-to-end cooperative driving framework. Equipped with a dual-view perception architecture, AURORA mitigates spatial and semantic discrepancies across ego and roadside viewpoints through a query-level Cross-View Query Alignment and Fusion (CQAF) module. Leveraging the resulting unified tokens, a LoRA-adapted VLM bridges semantic reasoning and generative trajectory planning. Extensive closed-loop evaluations on V2XBench demonstrate that AURORA achieves state-of-the-art performance in heavily occluded scenarios, with a Route Completion rate of 98.21% and a Driving Score of 76.02, while requiring low roadside communication bandwidth. Ultimately, this work pioneers an extensible V2X--VLM paradigm, paving the way for next-generation cooperative autonomous driving.