about
WAFT: Warping-Alone Field Transforms for Optical Flow (2025) (arxiv.org)
2 points by peter_d_sherman 37 days ago | hide | past | pdf | 1 comment on HN

In plain words: To track how each pixel moves between two video frames, it shifts one image by its motion guess and compares pixels directly, instead of building a grid comparing every pixel pair. It tops three tests, running 1.3 to 4.1 times faster with less memory.

Abstract · WAFT: Warping-Alone Field Transforms for Optical Flow

We introduce Warping-Alone Field Transforms (WAFT), a simple and effective method for optical flow. WAFT is similar to RAFT but replaces cost volume with high-resolution warping, achieving better accuracy with lower memory cost. This design challenges the conventional wisdom that constructing cost volumes is necessary for strong performance. WAFT is a simple and flexible meta-architecture with minimal inductive biases and reliance on custom designs. Compared with existing methods, WAFT ranks 1st on Spring, Sintel, and KITTI benchmarks, achieves the best zero-shot generalization on KITTI, while being 1.3-4.1x faster than existing methods that have competitive accuracy (e.g., 1.3x than Flowformer++, 4.1x than CCMR+). Code and model weights are available at \href{https://github.com/princeton-vl/WAFT}{https://github.com/princeton-vl/WAFT}.

Yihan Wang, Jia Deng
arXiv:2506.21526 · cs.CV · submitted Jun 26, 2025 · updated Feb 6, 2026
abstract · pdf · html

add comment on HN

>"Optical flow is a fundamental low-level vision task that estimates per-pixel 2D motion between video frames. It has many downstream applications, including 3D reconstruction and synthesis [...] action recognition [...] frame interpolation [...] and autonomous driving..."

[...]

>"WAFT is similar to RAFT but replaces cost volume with high-resolution warping, achieving better accuracy with lower memory cost."

[...]

>"Compared with existing methods, WAFT ranks 1st on Spring, Sintel, and KITTI benchmarks, achieves the best zero-shot generalization on KITTI, while being

1.3 − 4.1× faster than existing methods

that have competitive accuracy (e.g., 1.3× than Flowformer++, 4.1× than CCMR+)."