Urban perception is a complex cognitive process, yet existing street-view studies often overlook multidimensional contexts like spatial atmosphere and semantics. To bridge this gap, we propose Urban Perception CLIP (UP-CLIP), a specialized multimodal framework for extracting fine-grained perceptual semantics. Addressing the scarcity of high-quality subjective data, we developed a lexicon-guided, human-AI collaborative annotation paradigm to construct a domain-specific dataset of 15,000 image-text pairs. By fine-tuning CLIP on this dataset, we aligned street-view visual embeddings with multidimensional perception texts. Quantitative evaluations show that UP-CLIP outperforms pre-trained baselines, achieving a Recall@5 over 70% in bi-directional retrieval. Empirical applications in Wuhan demonstrate that UP-CLIP effectively: (1) maps the spatial distribution of ambivalent perceptions (e.g., ‘vibrant yet worn-out’) via text-to-image retrieval; and (2) generates city-wide sentiment distribution maps through image-to-text retrieval. Overall, UP-CLIP provides an automated, large-scale tool for urban perception, offering a reliable solution for researchers and practitioners to enhance urban environment quality.