论文
arXiv
SpatialIntelligence
Trajectory
Mobility
Agent
UrbanTraffic
中文标题
CoRe-MARL:基于循环多智能体强化学习的未知动态下协作再分配
English Title
CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning
Naimur Rahman Chowdhury, Shatabdi Sen Prapti, Md. Salehin Seyam, Limon Bin Hossain
发布时间
2026/9/16 21:25:05
来源类型
preprint
语言
en
摘要
中文对照

应急管理援助项目(如救灾物资分发)对于向受灾社区提供必要物资至关重要。然而,这些项目在由地方中心组成的去中心化网络中运行,各中心面临不确定的局部需求和供应动态,导致本地服务可用性不一致。在这些地方中心之间重新分配物资可以缓解这种不平衡,但各中心往往在信息有限且交通中断的情况下独立做出决策。本研究开发了 CoRe-MARL,这是一个协作多智能体强化学习(MARL)框架,通过构建去中心化部分可观测马尔可夫决策过程(Dec-POMDP)来实现。我们将每个中心视为一个智能体,学习一种再分配策略,以改善最差情况区域的服务水平并缩小区域间的服务差距,同时保障全网服务水平。我们引入了一种循环网络,用于捕捉无需直接观测即可演变的供需动态,而多智能体近端策略优化(MAPPO)则实现了集中式训练和分布式执行(CTDE)。我们在具有多样化轨迹的模拟环境中评估该框架,其中行动者和 MAPPO 评论家均无法观测到确切的动态。我们将循环 MAPPO 与循环独立 PPO(IPPO)以及仅本地启发式方法进行比较,发现 MAPPO 能够减少地方中心之间的服务差距,提升服务最差中心的服务水平,同时保持具有竞争力的全网服务水平。循环 MAPPO 在多样化的轨迹模式中表现出一致的性能,证明了其适应演变动态的能力。研究结果展示了协作学习在去中心化再分配中的能力。

English Original

Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs operate in a decentralized network of local centers that face uncertain local demand and supply dynamics, resulting in inconsistent avail- ability of local services. Redistribution of supplies among these local centers reduces these imbalances, but the centers often make decisions independently, with limited information and disrupted transportation. This study develops CoRe-MARL, a cooperative multi-agent reinforcement learning (MARL) framework, by formulating a decentralized partially observable Markov decision process (Dec-POMDP). We treat each center as an agent that learns a redistribution policy to improve the service in the worst-case region and reduce the service gap across regions while protecting network-wide service. We incorporate a recurrent network that captures evolving supply and demand dynamics without direct observation, while multi-agent proximal policy optimization (MAPPO) enables centralized training and decentralized execution (CTDE). We evaluate the framework in a simulated environment with diverse trajectories, where exact dynamics are not observed by actors and the MAPPO critic. We compare the recurrent MAPPO with the recurrent independent PPO (IPPO) and a local only heuristic, and find that MAPPO reduces the service gap across local centers and enhances service for the worst-served center while maintaining competitive network-wide service. The recurrent MAPPO also shows consistent performance across diverse trajectory patterns, demonstrating its ability to adapt to evolving dynamics. The findings demonstrate the capability of cooperative learning for decentralized redistribution and improving equitable service under uncertain and evolving dynamics.

我的阅读记录

正在加载阅读记录…

元数据
arXiv2609.18639v1
来源arXiv
类型论文
抽取状态raw
关键词
SpatialIntelligence
Trajectory
Mobility
Agent
UrbanTraffic
cs.LG
cs.AI