城市轨迹生成是交通仿真、城市规划与移动性分析的一项基础任务。然而,由于现有研究常采用不同的数据集、预处理流程、轨迹表示方法及评估指标,轨迹生成方法间的系统性比较仍十分困难。这种碎片化使得难以判断所报告的性能差异究竟源于生成机制本身,还是源于实验协议的不一致。为解决此问题,我们提出 CityTrajBench——一个面向城市尺度车辆轨迹生成的统一基准框架与协议。CityTrajBench 在统一设定下标准化了数据接入、轨迹归一化、特征构建、模型适配、地图感知后处理、模型选择及多层级评估。该基准支持多种异构生成器,包括统计基线模型、基于变分自编码器(VAE)、生成对抗网络(GAN)、扩散模型(diffusion)及流匹配(flow-matching)的模型,并在三个真实城市轨迹数据集上对其进行评估。基准评估维度涵盖全局空间真实性、行程级分布保真度、轨迹级几何相似性、条件移动一致性以及计算效率。实验揭示了不同模型族间的显著权衡:DiffTraj 在轨迹级几何保真度上表现最优;DiffRNTraj 在结构敏感的全局真实性方面具有竞争力;TrajFlow 则在真实性、质量、条件一致性与效率之间实现了良好平衡。与此同时,一个简单的马尔可夫基线模型在粗粒度行程统计与局部运动统计上仍具竞争力。这些发现表明,城市轨迹生成质量本质上是多目标的,不存在单一模型能在所有评估准则上全面占优,且 CityTrajBench 提供了一个可复现、可扩展、可比较的评估平台。
Urban trajectory generation is a fundamental task for transportation simulation, urban planning, and mobility analytics. However, systematic comparison across trajectory generation methods remains difficult because existing studies often rely on different datasets, preprocessing pipelines, trajectory representations, and evaluation metrics. This fragmentation makes it unclear whether reported performance differences arise from the generation mechanism itself or from inconsistent experimental protocols. To address this issue, we present CityTrajBench, a unified benchmark framework and protocol for city-scale vehicle trajectory generation. CityTrajBench standardizes data ingestion, trajectory normalization, feature construction, model adaptation, map-aware post-processing, model selection, and multi-level evaluation under a common setting. It supports heterogeneous generators, including statistical baselines, VAE-based, GAN-based, diffusion-based, and flow-matching-based models, and evaluates them on three real-world urban trajectory datasets. The benchmark measures global spatial realism, trip-level distribution fidelity, trajectory-level geometric similarity, conditional mobility consistency, and efficiency. Experiments reveal clear trade-offs across model families: DiffTraj is strongest on trajectory-level geometric fidelity, DiffRNTraj is competitive on structure-sensitive global realism, and TrajFlow provides a strong balance across realism, quality, conditional consistency, and efficiency. Meanwhile, a simple Markov baseline remains competitive on coarse-grained trip and local-movement statistics. These findings show that urban trajectory generation quality is inherently multi-objective, that no single model dominates all criteria equally, and that CityTrajBench provides a reproducible benchmark protocol and testbed for future research on urban mobility generation.