高质量的土地利用数据集对于推动土地利用分类与识别研究至关重要。然而,土地利用的复杂性与空间异质性给数据集构建带来了挑战。为应对这些问题,我们提出了MSLU-100K,一个涵盖81个中国城市超过10万个不规则地块样本的多源土地利用数据集。该数据集基于人机协同框架构建,融合遥感数据与兴趣点(POI)数据,将地块划分为7类主要土地利用类型和28类次要土地利用类型。采用一种新颖的多层级分类方法,结合人工标注与深度学习技术,确保数据在六个质量等级上的高可靠性。数据集中超过57%的样本属于高质量级别(4级和5级),显著提升了分类性能。该数据集为土地利用识别、城市规划及空间研究提供了坚实资源。
High-quality land use datasets are essential for advancing research in land use classification and recognition. However, the complexity and spatial heterogeneity of land use create challenges in dataset construction. To address these issues, we present MSLU-100K, a multi-source land use dataset encompassing over 100,000 irregular parcel samples from 81 Chinese cities. Constructed using a human-computer collaboration framework, this dataset integrates remote sensing and POI (Point of Interest) data, categorizing parcels into 7 primary and 28 secondary land use types. A novel multi-level classification approach combines manual labeling and deep learning, ensuring high data quality across six quality levels. Over 57% of the dataset comprises high-quality samples (Levels 4 and 5), which significantly enhance classification performance. The dataset provides a robust resource for land use recognition, urban planning, and spatial research.