访视熵(visitation entropy)即个体在不同地点间访问分布的香农熵(Shannon entropy),是人类移动性研究中广泛使用的度量指标。然而,其广泛应用依赖于若干鲜少被明示的假设:熵需定义在一组固定的状态集上,且其实证估计要求大量、充分采样的观测数据。当熵在此设定之外被使用时所引发的局限性,已在统计物理、生态学和密码学等领域被记录并深入探讨;但其对移动性研究的影响尚不明确。本文利用合成轨迹与实证轨迹,系统考察访视熵在刻画人类移动性方面的优势与缺陷。我们发现,对于个体访问的一系列地点序列,访视熵主要反映所访问的独特地点数量与序列长度,二者共同解释了实证数据中访视熵方差的90.7%。我们还发现,较短的序列会系统性地低估熵值,表明有限样本偏差(finite-sample bias)是一项关键局限。因此,基于访视熵对不同群体的比较应持审慎态度。我们发现,在控制序列长度与独特地点数量后,性别之间、通勤者与非通勤者之间、城市居民与农村居民之间的表观熵差异显著缩小,甚至发生逆转。最后,我们指出,可通过控制序列长度与独特地点数量,并辅以结构化网络度量(structural network measures)来缓解上述问题;后者能提供比单纯熵更细致的洞察,揭示移动性组织方式中熵本身无法捕捉的方面。
Visitation entropy, the Shannon entropy of an individual's distribution of visits across locations, is a widely used metric in the human mobility literature. Yet, its widespread use rests on assumptions that are rarely made explicit: entropy is defined over a fixed set of states, and estimating it empirically requires abundant, well-sampled observations. The limitations that arise when entropy is used outside this setting have been documented and explored in fields such as statistical physics, ecology, and cryptography. The implications for mobility studies, however, remain unclear. Here, we leverage synthetic and empirical trajectories to systematically examine the strengths and weaknesses of visitation entropy as a measure for characterizing human mobility. We show that, for a sequence of locations visited by an individual, the visitation entropy primarily reflects the number of unique locations visited and sequence length, which together explain 90.7% of its variance in empirical data. We also show that shorter sequences systematically lead to an underestimation of entropy, showing finite-sample bias to be a key limitation. As a result, comparisons of groups based on visitation entropy should be treated with caution. We find that apparent entropy gaps between genders, commuters and non-commuters, and urban and rural residents are reduced, or reversed, once sequence length and the number of unique locations are taken into account. Finally, we show that these issues can be addressed by controlling for sequence length and the number of unique locations, and by complementing entropy with structural network measures, which provide more nuanced insight into how mobility is organized beyond the aspects that entropy alone captures.