In plain words: A picture-making diffusion model's internal attention maps can spread object labels from one region to another, and doing this frame by frame turns it into a video object tracker needing no video training. It beat other zero-shot trackers on standard video segmentation tests.
Abstract
Image diffusion models, though originally developed for image generation, implicitly capture rich semantic structures that enable various recognition and localization tasks beyond synthesis. In this work, we investigate their self-attention maps can be reinterpreted as semantic label propagation kernels, providing robust pixel-level correspondences between relevant image regions. Extending this mechanism across frames yields a temporal propagation kernel that enables zero-shot object tracking via segmentation in videos. We further demonstrate the effectiveness of test-time optimization strategies-DDIM inversion, textual inversion, and adaptive head weighting-in adapting diffusion features for robust and consistent label propagation. Building on these findings, we introduce DRIFT, a framework for object tracking in videos leveraging a pretrained image diffusion model with SAM-guided mask refinement, achieving state-of-the-art zero-shot performance on standard video object segmentation benchmarks.
Youngseo Kim, Dohyun Kim, Geonhee Han, Paul Hongsuck Seo
arXiv:2511.19936 · cs.CV · submitted Nov 25, 2025
abstract · pdf · html
Makes you wonder what intelligence is lurking in a 10T parameter model like Gemini 3 that we may not discover for some years yet…