about
SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image Editing (arxiv.org)
2 points by PaulHoule 320 days ago | hide | past | pdf | discuss on HN

In plain words: This editor skips mapping the photo back into the model's hidden space: it splits the target description into parts, builds a path for each, and blends them to cut overlap. It edits more faithfully and separates changed traits better than other editors that skip this step. Count check: S1 = 30, S2 = 16 → 46. Over. Fix: drop "the photo" → "a photo"? same. Drop "and" → "then"? same. Remove "the model's hidden space" → "its hidden space" (saves 1) = 29; total 45. Still not under 45. Remove "mapping" → "pushing"? same length. Remove "the target description" → "the target" (saves 2) = 28; total 44. ✓ Final: "This editor skips mapping the photo back into the model's hidden space: it splits the target into parts, builds a path for each, and blends them to cut overlap. It edits more faithfully and separates changed traits better than other editors that skip this step." Hmm, "the target" is a bit vague but acceptable — target prompt. Actually "the target prompt" = 2 words vs "the target description" = 2 words... "target description" is 2 words, "target prompt" is 2 words. Same. Fine, use "target description" and cut elsewhere: drop "the model's hidden space" → "its hidden space" saves 1 → 29, total 45. Not under. Drop "a path for each" → "a path each"? no. Drop "blends them to cut overlap" → "blends them, cutting overlap" saves 1 → 28 with "target description". Total 44. ✓ "This editor skips mapping the photo back into the model's hidden space: it splits the target description into parts, builds a path for each, and blends them, cutting overlap. It edits more faithfully and separates changed traits better than other editors that skip this step." Hmm "cutting overlap" — the weights suppress redundancy. Fine. Actually simpler: "blends them with weights that cut overlap" is clearer but longer. Let's keep "blends them, cutting overlap" — acceptable. Final answer below.

Abstract

Rectified flow models have become a de facto standard in image generation due to their stable sampling trajectories and high-fidelity outputs. Despite their strong generative capabilities, they face critical limitations in image editing tasks: inaccurate inversion processes for mapping real images back into the latent space, and gradient entanglement issues during editing often result in outputs that do not faithfully reflect the target prompt. Recent efforts have attempted to directly map source and target distributions via ODE-based approaches without inversion; however,these methods still yield suboptimal editing quality. In this work, we propose a flow decomposition-and-aggregation framework built upon an inversion-free formulation to address these limitations. Specifically, we semantically decompose the target prompt into multiple sub-prompts, compute an independent flow for each, and aggregate them to form a unified editing trajectory. While we empirically observe that decomposing the original flow enhances diversity in the target space, generating semantically aligned outputs still requires consistent guidance toward the full target prompt. To this end, we design a projection and soft-aggregation mechanism for flow, inspired by gradient conflict resolution in multi-task learning. This approach adaptively weights the sub-target velocity fields, suppressing semantic redundancy while emphasizing distinct directions, thereby preserving both diversity and consistency in the final edited output. Experimental results demonstrate that our method outperforms existing zero-shot editing approaches in terms of semantic fidelity and attribute disentanglement. The code is available at https://github.com/Harvard-AI-and-Robotics-Lab/SplitFlow.

Sung-Hoon Yoon, Minghan Li, Gaspard Beaudouin, Congcong Wen, Muhammad Rafay Azhar, Mengyu Wang
arXiv:2510.25970 · cs.CV · submitted Oct 29, 2025
abstract · pdf · html · Camera-ready version for NeurIPS 2025, 10 pages (main paper)

add comment on HN