面向自动驾驶的最先进视觉-语言-动作(VLA)模型面临关键局限:参数规模过大、高分辨率图像处理效率低下以及缺乏时序记忆。我们提出 Fast and EffectIVE VLA(FIVE-VLA),通过两项核心贡献解决上述问题。首先,采用高效的视觉编码器处理高分辨率(448 × 896)图像,仅生成 98 个 token,比现有方法减少超过 5 倍,并完全绕过文本生成以实现单步轨迹预测。其次,提出循环动作记忆(Recurrent Action Memory, RAM),这是一种轻量级模块,利用先前动作 token 对动作预测进行条件约束,为超车和紧急制动等机动操作提供关键的时序上下文。FIVE-VLA 仅有 6.41 亿参数,在具有挑战性的 Bench2Drive 闭环驾驶基准测试中,相比此前最先进的 VLA 模型,无交通违规规则完成的路线比例高出约 10%。在大规模真实世界 NVIDIA Physical AI AV 数据集上的非反应式开环仿真显示,其碰撞违规率分别比 SimLingo 低 10.2%(单视图设置)和 7.7%(四视图设置)。此外,FIVE-VLA 在 A100 GPU 上运行速度约为 30 fps,在 T4 GPU(作为边缘设备的代理)上约为 4 fps,相比先前方法实现了 8 至 30 倍的加速。
State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes $\sim$10% more routes without traffic rule infractions than the previous state-of-the-art VLA on the challenging Bench2Drive closed-loop driving benchmark. Non-reactive open-loop simulation on the large-scale real-world NVIDIA Physical AI AV dataset shows 10.2% and 7.7% lower collision-violation rates than SimLingo in single- and four-view settings, respectively. Additionally, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an 8-30$\times$ speedup over previous methods.