在信号控制交叉口的困境区(dilemma zone)内,驾驶员决策关乎行车安全:车辆在黄灯亮起时需在有限的时间与距离约束下决定停车或通行。对停车/通行决策及其决策时刻的准确预测,对自适应信号控制、高级驾驶辅助系统及以人为中心的智能交通应用具有重要意义。然而,困境区行为具有显著的驾驶员个体依赖性:相似的接近轨迹可能因风险偏好、制动习惯及决策阈值的差异而在不同驾驶员间产生不同决策。现有个性化模型多依赖人工设计的标量描述符,虽具一定效用,但对个体行为的刻画较为有限。本文提出VISTA-DZ——一种基于语义画像条件化的个性化停车/通行及决策时刻预测框架。该方法将历史轨迹转化为视觉表征,由视觉-语言模型解析生成行为画像,并编码为语义嵌入以调节双输出预测网络。最终模型融合双向GRU编码器、驾驶员条件化多头交叉注意力机制以及特征级线性调制(Feature-wise Linear Modulation),实现时序证据选择与特征自适应。在SDZ数据集及新采集的FDZ数据集上的实验表明,VISTA-DZ优于仅使用轨迹信息及基于人工特征的个性化基线模型,在领域内仿真中达到93.26%的准确率,在20名预留仿真驾驶员上的平均准确率为90.22%。跨域结果进一步表明,该方法具备可行的零样本仿真到实车迁移能力,且当仿真数据与少量实地数据结合时,可提升真实场景泛化性能。
Driver decision making in the dilemma zone at signalized intersections is safety critical, as vehicles approaching a yellow signal must decide whether to stop or proceed within limited time and distance margins. Accurate prediction of both stop-go decisions and decision timing is important for adaptive signal control, advanced driver assistance systems, and human-centered intelligent transportation applications. However, dilemma zone behavior is strongly driver dependent. Similar approach trajectories may lead to different decisions across drivers because of differences in risk preference, braking habit, and decision threshold. Existing personalized models often rely on handcrafted scalar descriptors, which provide useful but limited summaries of individual behavior. This paper proposes VISTA-DZ, a semantic-profile-conditioned framework for personalized stop-go and decision-time prediction. Historical trajectories are converted into visual representations, interpreted by a vision-language model to generate behavioral profiles, and encoded as semantic embeddings to condition a dual-output prediction network. The final model combines a bidirectional GRU encoder, driver-conditioned multi-head cross-attention, and Feature-wise Linear Modulation for temporal evidence selection and feature adaptation. Experiments on the SDZ dataset and a newly collected FDZ dataset show that VISTA-DZ outperforms trajectory-only and handcrafted personalization baselines, achieving 93.26% in-domain simulation accuracy and 90.22% mean accuracy across 20 held-out simulation drivers. Cross-domain results further show feasible zero-shot simulation-to-real transfer and better real-world generalization when simulation data are combined with limited field data.