论文
arXiv
SpatialIntelligence
Trajectory
Mobility
GeoSimulation
中文标题
FlexComposer:面向图像到动态影像的统一视频合成框架,支持灵活轨迹控制
English Title
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
Songchun Zhang, Sitong Guo, Xianghao Kong, Pengwei Liu, Yuwei Guo, Lvmin Zhang, Anyi Rao
发布时间
2026/8/1 00:59:46
来源类型
preprint
语言
en
摘要
中文对照

生成式视频合成——即将外部素材无缝嵌入现有视频序列——在内容创作与视觉特效中至关重要。然而,现有方法普遍存在控制性与保真度之间的权衡:或从静态图像中幻化运动,导致预动画素材的动态特性无法保留;或缺乏细粒度空间控制,难以沿用户定义轨迹精确放置素材。我们提出 FlexComposer,一种将视频合成标准化为轨迹引导条件生成任务的统一框架,支持静态图像与动态影像的无缝融合。本方法包含三项核心设计:(1)统一规范前景表示(Unified Canonical Foreground Representation),将物体固有运动与其全局位移解耦,将异构输入标准化为稳定、居中的潜在空间;(2)空间感知潜在特征注入策略(Spatial-Aware Latent Injection),利用变分自编码器(VAE)潜在空间的平移等变性,通过无参数机制将规范特征映射至目标轨迹;(3)混合数据集与合成到真实渐进式训练范式(Hybrid Dataset and Synthetic-to-Real Curriculum),协同整合程序化仿真、真实电影镜头及生成数据,隐式学习符合物理规律的光照与阴影协调。该统一设计可处理从产品照片到动态主体的多样化输入,在无需显式三维重建或辅助可学习适配器的前提下,实现高保真运动控制与环境融合。大量实验表明,FlexComposer 在视觉质量、时序一致性与轨迹贴合度方面均优于当前最优方法。

English Original

Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and visual effects. However, existing approaches suffer from a control-fidelity trade-off: they either hallucinate motion from static images, failing to preserve the dynamics of pre-animated assets, or lack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage. Our approach introduces three key designs: (1) a Unified Canonical Foreground Representation that decouples an object's intrinsic motion from its global displacement, standardizing heterogeneous inputs into a stabilized, centered latent space; (2) a Spatial-Aware Latent Injection strategy that exploits the translation equivariance of VAE latent spaces to transport canonical features onto target trajectories via a parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cinematic footage, and generative data to implicitly learn physically plausible illumination and shadow harmonization. This unified design handles diverse inputs from product photos to dynamic subjects achieving high-fidelity motion control and environmental integration without the need for explicit 3D reconstruction or auxiliary learnable adapters. Extensive experiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory adherence.

元数据
arXiv2607.29627v1
来源arXiv
类型论文
抽取状态raw
关键词
SpatialIntelligence
Trajectory
Mobility
GeoSimulation
cs.CV