城市交通拥堵是一个日益严重的全球性问题,显著加剧了通勤时间延长与环境污染。传统交通信号控制系统往往难以适应动态变化的交通状况。自适应交通信号控制可在不改变道路基础设施的前提下改善城市交通。深度强化学习(DRL)在此任务中展现出优异性能,但现有基于延误或排队长度的奖励函数常导致短视或不稳定的策略。本文提出一种基于动量的奖励函数(MBRF),旨在鼓励车辆持续通行,而非仅对拥堵施加惩罚。该方法在SUMO(Simulation of Urban MObility)仿真平台中进行评估,采用等待时间、排队长度、通行量及CO2排放等标准交通指标。结果表明,相较于基于延误或排队长度的奖励函数以及经典控制器(如Max Pressure和LQF),所提奖励函数可实现更优的通行量-排放权衡,并展现出更稳定的训练行为。
Urban traffic congestion is a growing global issue contributing significantly to long commute times and environmental pollution. Traditional traffic signal control systems often fail to adapt to dynamic traffic conditions. Adaptive traffic signal control can improve urban traffic without changing road infrastructure. Deep Reinforcement Learning (DRL) has shown strong performance for this task, but existing delay and queue-based rewards often produce short-sighted or unstable policies. This paper proposes a Momentum-Based Reward Function (MBRF) that encourages vehicles to keep moving rather than penalizing congestion alone. The method is evaluated in SUMO (Simulation of Urban MObility) using standard traffic metrics such as waiting time, queue length, throughput, and CO2 emissions. Results show that the proposed reward produces better throughput-emission trade-offs and more stable learning behavior than delay or queue-based rewards, as well as classical controllers such as Max Pressure and LQF.