共享微出行研究通常基于行程建模与用户数据。传统上,共享微出行系统的用户细分方法是先将行程层面的观测值聚合为用户层面的汇总统计量,再应用聚类技术;此类聚合可能掩盖行程层面的变异性,并在将结果解释为个体记录特征时导致生态谬误。本文提出一种针对多元分类计数数据的贝叶斯有限混合模型,该模型直接基于重复的行程层面观测对用户进行聚类,同时完整保留个体出行行为的分类结构。该方法聚焦于从高维分类行程行为中识别异质性出行用户,并在聚类分配中纳入不确定性。用户是探索潜在聚类模式的基本分析单元。模型以乘积多项式似然函数表征每位用户,并引入潜在聚类隶属关系。本方法以意大利威尼斯市政府提供的为期一年的共享单车与共享电单车行程记录数据为例进行说明,该数据集包含逾11,000名高频用户的22万余条行程记录。分析识别出八类不同的潜在出行画像,分别对应本地化出行、通勤导向型、游客导向型、中心城区内出行及跨区域出行等行为模式。所提出的框架具备灵活性与计算可扩展性,适用于对重复性分类观测进行聚类,并可直接推广至其他大规模行为与交通数据集。
The study on shared micro-mobility is based on trip modeling and user data. User segmentation in shared micromobility systems is traditionally studied by aggregating trip-level observations into user-specific summary measures before applying clustering techniques. Such aggregation can obscure trip-level variability and lead to ecological fallacies if results are interpreted as applying to individual records. We propose a Bayesian finite mixture model for multivariate categorical count data that clusters users directly from repeated trip-level observations while preserving the full categorical structure of individual travel behavior. This approach focuses on identifying heterogeneous mobility users from high-dimensional categorical trip behavior while accounting for uncertainty in cluster assignments. Users are the fundamental unit of analysis for exploring latent cluster patterns. The model represents each user with a product-multinomial likelihood with latent cluster membership. The methodology is illustrated using a one-year trip record of shared bikes and e-bikes from the Municipality of Venice, Italy, comprising over 220,000 trips made by more than 11,000 recurrent users. The analysis identifies eight distinct latent mobility profiles corresponding to localized, commuter-oriented, tourist-oriented, central, and inter-zonal travel behaviors. The proposed framework provides a flexible and computationally scalable approach for clustering repeated categorical observations and is readily applicable to other large-scale behavioral and transportation datasets.