智慧城市快速发展日益依赖于轨迹数据挖掘,但公共出行数据集中往往缺乏对弱势人口群体(尤其是老年人)的充分表征。这种表征不足可能在出行建模及下游城市规划中引入系统性偏差。本研究利用2016–2020年纽约市Citi Bike系统数据中的泽西城子集,以合成轨迹生成为案例,定量考察弱势亚群体出行特征缺失对出行建模的影响。分析表明,老年骑行者展现出结构上显著不同的出行特征:活动空间更局域化(958米 vs. 年轻骑行者的1189米)、出行熵更低(1.82 vs. 4.15)、非高峰时段的时间分布呈不对称性。为验证仅依赖多数群体主导的训练数据会导致合成结果产生偏差,我们进一步在三种人口统计学训练设定下——全量人群、仅年轻骑行者、仅老年骑行者——分别评估了一阶马尔可夫链模型与经QLoRA微调的Qwen3-4B大语言模型。结果表明,基于多数群体主导数据训练的模型系统性地误表征老年群体的出行行为,尤其在空间出行指标上偏差显著:全量人群训练的马尔可夫模型对老年骑行者步长和停留时长的估计分别高估4.5%和8.9%,而仅使用老年骑行者数据训练的模型在多数指标上误差显著降低。马尔可夫模型与大语言模型框架的对比还显示,在人口统计数据有限的情况下,更高能力的模型未必能提升亚群体层面的表征保真度。上述发现凸显了人口统计学代表性对出行建模及其下游应用的重要性。
The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the elderly, are often sparsely represented in public mobility datasets. This underrepresentation can introduce systematic bias into mobility modeling and downstream urban planning. Using the 2016-2020 Jersey City subset of the Citi Bike System Data, this study quantitatively examines how the absence of underrepresented subgroups' mobility signatures affects mobility modeling, using synthetic trajectory generation as a case study. The analysis reveals that elderly riders exhibit a structurally distinct mobility signature, including localized activity spaces (958 m vs. 1,189 m for young riders), lower mobility entropy (1.82 vs. 4.15), and asymmetric off-peak temporal patterns. To demonstrate that relying on majority-dominated training data yields biased synthetic outcomes, we further evaluate both a first-order Markov chain and a Qwen3-4B model fine-tuned with QLoRA across three demographic training settings: the full population, young riders only, and elderly riders only. Results show that models trained on majority-dominated populations systematically misrepresent elderly mobility behavior, particularly for spatial mobility metrics. The Markov model trained on the full population overestimates elderly step length by 4.5% and dwell time by 8.9%, whereas the elderly-specific model achieves substantially lower errors across most metrics. Comparisons between the Markov and LLM-based frameworks further show that higher-capability models do not necessarily improve subgroup-level fidelity under limited demographic data. These findings underscore the importance of demographic representation in mobility modeling and its downstream applications for underrepresented populations.