论文
arXiv
LLM
Multimodal
中文标题
多模态城市感知中的角色提示:描述性收敛与解释性变异
English Title
Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation
Neemias da Silva, Matt Ratto, Myriam Delgado, Rodrigo Minetto, Daniel Silver, Thiago H Silva
发布时间
2026/5/28 04:11:42
来源类型
preprint
语言
en
摘要
中文对照

本研究考察角色提示(persona prompting)如何影响两个多模态大语言模型(Qwen3-VL-8B 和 Gemma-4-E4B-it)在城市感知任务中生成的语言,该任务为探究共享视觉证据下主观解释的差异提供了典型场景。我们将模型输出划分为三个功能层级:描述性基础层(图像描述)、中间语义层(感知标签)和解释性框架层(理由陈述)。基于每模型约 60,000 条经角色条件化标注的数据,我们发现图像描述在不同角色档案间高度收敛,仅在属性关联维度上呈现微小差异;而理由陈述的变异程度显著更高:经济地位在两个模型中均引发最大差异,政治取向与人格特质亦具显著影响。配对图像级比较进一步证实,上述三类属性所导致的理由陈述差异均大于其图像描述差异。对于感知标签,具有相同属性水平的角色生成的标签集合相似度高于属性水平不同的角色,其中经济地位所导致的分离度最大。探索性主题分析还揭示了角色特异性的评价侧重。跨模型分析显示,所有三类输出中角色档案对之间的相似性模式均呈强相关性,但理由陈述的一致性最低。总体而言,角色提示对解释性框架的影响强于对描述性基础的影响。

English Original

This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional levels: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations per model from Qwen3-VL-8B and Gemma-4-E4B-it, we find that captions converge strongly across persona profiles and show only small attribute-associated differences. Justifications vary substantially more: economic status produces the largest difference in both models, with political orientation and personality also prominent. Paired image-level comparisons confirm larger justification than caption differences for these three attributes. For perception tags, personas sharing the same attribute level produce more similar tag sets than personas with different attribute levels, with the largest separation observed for economic status. Exploratory topic analysis further reveals persona-specific evaluative emphasis. Across models, profile-pair similarity patterns are strongly correlated for all three output types, although agreement is lowest for justifications. Overall, persona prompting affects interpretive framing more strongly than descriptive grounding.

我的阅读记录

正在加载阅读记录…

元数据
arXiv2605.29064v2
来源arXiv
类型论文
抽取状态raw
关键词
LLM
Multimodal
cs.CL
cs.CV
cs.HC
cs.MA