• 首页
  • 期刊简介
  • 编委会
  • 投稿指南
  • 收录情况
  • 杂志订阅
  • 联系我们
引用本文:薛梦娇,叶宝林,张一嘉.混合动作空间下的深度强化学习交通信号控制方法[J].软件工程,2026,29(6):61-66.【点击复制】
【打印本页】   【下载PDF全文】   【查看/发表评论】  【下载PDF阅读器】  
←前一篇|后一篇→ 过刊浏览
分享到: 微信 更多
混合动作空间下的深度强化学习交通信号控制方法
薛梦娇,叶宝林,张一嘉
(浙江理工大学信息科学与工程学院, 浙江 杭州 310018)
942636599@ qq.com; yebaolin@ zjxu.edu.cn; waiting@ zstu.edu.cn
摘 要: 现有基于深度强化学习的交通信号控制方法需提前定义绿灯相位的时间或顺序,在应对复杂交通流量时缺乏灵活性。为此,提出了一种混合动作空间下的深度强化学习方法HA-TD3。该方法将交通信号控制建模为混合动作空间下的马尔可夫决策过程,结合双延迟深度确定性策略梯度(TD3)算法同步优化绿灯相位选择和绿灯相位时间。首先,构建绿灯相位和持续时间之间的依赖关系,引入条件变分自编码器;其次,为挖掘车流数据的空间相关性,设计了离散交通状态编码方法和空间信息提取模块;最后,基于微观交通仿真软件SUMO的仿真测试结果表明,与基准深度强化学习方法相比,所提方法在800、1000、1200辆/h流量下车辆的排队长度和等待时间分别减少了20.4%、21.6%和37.1%以及13.3%、12.8%和23.7%。
关键词: 深度强化学习  交通信号控制  混合动作空间  条件变分自编码器
中图分类号: TP181    文献标识码: A
基金项目: 浙江省自然科学基金项目资助(LTGS23F030002);嘉兴市应用性基础研究项目(2023AY11034);工业控制技术国家重点实验室开放课题(ICT2022B52)
Deep Reinforcement Learning for Traffic Signal Control in Hybrid Action Spac
XUE Mengjiao, YE Baolin, ZHANG Yijia
(School of Information Science and Engineering, Zhejiang Sc-i Tech University, Hangzhou 310018, China)
942636599@ qq.com; yebaolin@ zjxu.edu.cn; waiting@ zstu.edu.cn
Abstract: Existing traffic signal control methods based on deep reinforcement learning require the green light phase time or sequence to be predetermined in advance, which lack flexibility in dealing with complex traffic flow. Therefore, this paper proposes a hybrid action space deep reinforcement learning method, HA-TD3. The method models traffic signal control as a mixed action space Markov decision process and optimizes the selection and duration of green light phases simultaneously by combining the Twin Delayed Deep Deterministic policy gradient (TD3) algorithm. First, the dependence between green light phases and their durations is constructed, and a conditional variational autoencoder is introduced. Second, a discrete traffic state encoding method and a spatial information extraction module are designed to mine the spatial correlation of traffic flow data. Finally, the simulation test results based on the microscopic traffic simulation software SUMO show that compared with the benchmark deep reinforcement learning method, the proposed method reduces the vehicle queue length and waiting time by 20. 4% , 21. 6% , and 37. 1% respectively, and reduces them by 13.3% , 12.8% , and 23.7% in the three traffic flow scenarios.
Keywords: deep reinforcement learning  traffic signal control  hybrid action space  conditional variational autoencoder


版权所有:软件工程杂志社
地址:辽宁省沈阳市浑南区创新路195号 邮政编码:110169
电话:0411-84767887 传真:0411-84835089 Email:semagazine@neusoft.edu.cn
备案号:辽ICP备17007376号-1
技术支持:北京勤云科技发展有限公司

用微信扫一扫

用微信扫一扫