论文
International Journal of Digital Earth
PublisherJournal
UrbanComputing
中文标题
从像素到感知:面向语义城市感知的视觉-语言模型微调
English Title
From pixels to perception: fine-tuning vision-language models for semantic urban perception
Guanjie Huang Xifa Song Hongzan Jiao a Department of Urban Planning, School of Urban Design, Wuhan University, Wuhan, People's Republic of China b Aerospace Information Research Institute, Henan Academy of Sciencess, Henan, People's Republic of China
发布时间
2026/5/25 13:02:55
来源类型
journal
语言
en
摘要

Urban perception is a complex cognitive process, yet existing street-view studies often overlook multidimensional contexts like spatial atmosphere and semantics. To bridge this gap, we propose Urban Perception CLIP (UP-CLIP), a specialized multimodal framework for extracting fine-grained perceptual semantics. Addressing the scarcity of high-quality subjective data, we developed a lexicon-guided, human-AI collaborative annotation paradigm to construct a domain-specific dataset of 15,000 image-text pairs. By fine-tuning CLIP on this dataset, we aligned street-view visual embeddings with multidimensional perception texts. Quantitative evaluations show that UP-CLIP outperforms pre-trained baselines, achieving a Recall@5 over 70% in bi-directional retrieval. Empirical applications in Wuhan demonstrate that UP-CLIP effectively: (1) maps the spatial distribution of ambivalent perceptions (e.g., ‘vibrant yet worn-out’) via text-to-image retrieval; and (2) generates city-wide sentiment distribution maps through image-to-text retrieval. Overall, UP-CLIP provides an automated, large-scale tool for urban perception, offering a reliable solution for researchers and practitioners to enhance urban environment quality.

我的阅读记录

正在加载阅读记录…

元数据
来源International Journal of Digital Earth
类型论文
抽取状态raw
关键词
PublisherJournal
UrbanComputing