城市交通拥堵在以汽车为主导的城市中仍是长期存在的挑战,带来显著的经济与社会成本。交通信号系统正日益作为网络化信息物理系统组件部署于智慧城市基础设施中,分布式感知与边缘智能支撑了自适应交通管理。本文研究强化学习(RL)作为一种边缘智能方法,用于科威特某信号控制城市交叉口的自适应交通信号运行。本文设计了一种基于近端策略优化(PPO)的控制器,仅依据本地观测到的交通状态动态分配绿灯时长,无需未来交通需求信息或中心化协调。该控制器在基于科威特真实小时级交通流量数据构建的高保真仿真环境中进行评估,并以平均车辆延误、排队长度和排放量为性能指标,与传统定时控制及代表当前工程实践的车辆感应式控制器进行对比。在基准工况下,所提控制器相较定时控制降低平均车辆延误46%,相较感应式控制降低34%,同时使单车CO2排放量减少约23%。上述性能优势在交通需求波动±15%条件下依然保持,可从工作日交通模式泛化至周末模式,并经奖励函数消融实验验证;五组随机种子实验结果方差较低,证实其统计可靠性。本研究结果表明,基于学习的边缘交通信号控制具备实际应用价值,可作为物联网赋能的智慧城市交通系统的构建模块,并为全面互联的车联网(IoV)系统提供可部署的前期技术路径。
Urban traffic congestion remains a persistent challenge in car-dependent cities, imposing significant economic and societal costs. Traffic signal systems are increasingly deployed as networked cyber-physical components within smart-city infrastructures, where distributed sensing and edge intelligence enable adaptive traffic management. This paper investigates reinforcement learning (RL) as an edge-intelligent approach for adaptive traffic signal operation at a signalized urban intersection in Kuwait. A Proximal Policy Optimization (PPO)-based controller is developed to dynamically allocate green-phase durations using locally observed traffic states, without relying on future demand information or centralized coordination. The controller is evaluated in a realistic simulation environment informed by real-world hourly traffic volume data from Kuwait, and is compared against both conventional fixed-time control and a vehicle-actuated controller representing the current state of practice, using average vehicle delay, queue length, and emissions as performance metrics. Under nominal conditions, the proposed controller reduces average vehicle delay by 46% relative to fixed-time control and 34% relative to actuated control, while also lowering per-vehicle CO2 emissions by approximately 23%. These performance gains persist under demand perturbations of +/-15%, generalize from weekday to weekend traffic patterns, and are corroborated by a reward function ablation; low variance across five random seeds confirms their statistical reliability. These findings demonstrate the practicality of learning-based edge traffic signal control as a building block for IoT-enabled smart-city transportation systems, and as a deployable precursor toward fully connected, Internet of Vehicles (IoV)-based urban mobility.