Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model
arXiv 2026
Abstract
We show that a strong instruction-based image editing model can edit videos by operating directly on video-VAE latents. Starting from Qwen-Image-Edit, we arrange the latent frames of a Wan 2.1 video VAE as tiles of one large virtual image, reuse the editor's image positional encoding for every tile, and bridge the two latent spaces with a pair of lightweight input/output projections warm-started from the editor's own patchify and unpatchify layers. The system is fine-tuned on the public Ditto-1M editing triplets, with a few denoising steps of Wan 2.2 as an optional temporal enhancer. Our results suggest that per-frame video latents remain close enough to the image domain that mature image editing priors transfer with minimal adaptation.












