论文
International Journal of Digital Earth
PublisherJournal
Platform
中文标题
基于视觉Transformer与多任务学习的栅格地图建筑泛化框架
English Title
A visual transformer and multi-task learning framework for building generalisation from raster maps
Junbo Yu Guanglei Pan Yu Feng Yi Xiao Aji Gao Li Liu Huafei Yu Tinghua Ai Dong Ren a College of Computer and Information Technology, China Three Gorges University, Yichang, Hubei, People's Republic of China b Key Laboratory of Digital Mapping and Land Information Application, Minisitry of Natural Resources, Wuhan, Hubei, People's Republic of China c Hubei Key Laboratory of Intelligent Vision Based Monitoring for Hydroelectric Engineering, China Three Gorges University, Yichang, Hubei, People's Republic of China d i3mainz - Institute for Spatial Information and Surveying Technology, University of Applied Sciences Mainz, Mainz, Germany e School of Computer and Software, Shenzhen University of Information Technology, Shenzhen, People's Republic of China f National Engineering Research Center for Digital Construction and Evaluation of Urban Rail Transit, Tianjin, People's Republic of China g School of Resource and Environmental Sciences, Wuhan University, Wuhan, Hubei, People's Republic of China
发布时间
2026/8/31 12:08:47
来源类型
journal
语言
en
摘要

Building generalisation transforms detailed large-scale building representations into suitable simplified forms for smaller-scale urban topographic maps. Existing deep learning-based approaches for raster building maps are largely built on convolutional neural networks (CNNs) or generative adversarial networks (GANs). However, their local inductive biases or limited global spatial awareness often cause geometric distortions in buildings. Although large-vision foundation models excel in holistic modeling, their application has focused mainly on building extraction. To address these issues, SAM-BG, a Building Generalisation framework that integrates the Segment Anything Model with multitask learning, was proposed. Specifically, low-rank adaptation was employed to efficiently fine-tune the vision transformer (ViT) encoder and transfer global representation capability to the cartographic domain. Subsequently, a geometry-aware multitask architecture with explicit boundary supervision was introduced to alleviate contour distortion. Moreover, a Boundary Feature Enhancement Module (BFEM) that uses dual-source-guided gating and multi-scale feature pyramids was designed to reinforce boundary responses. Multi-scale experiments on synthetic and real-world datasets showed that SAM-BG outperformed benchmark methods on boundary-sensitive metrics of buildings. This study presents a viable end-to-end solution for building generalisation and validates the potential of vision foundation models within domain-specific context of cartography.

我的阅读记录

正在加载阅读记录…

元数据
来源International Journal of Digital Earth
类型论文
抽取状态raw
关键词
PublisherJournal
Platform