Abstract
Regional image editing requires more than following an instruction: the edit must stay within the intended region, preserve surrounding content, and blend naturally at the boundary. MaskFlow addresses these challenges with a unified training and inference framework. Mask-guided Siamese Probability Paths coordinate foreground generation with background preservation, while a mask-aware flow-matching objective focuses learning on the editable region. Soft-Poisson De-seaming refines the predicted vector field during both training and sampling for smooth boundary transitions. We also introduce MaskEdit-Benchmark, covering general scenes and infographics with concise instructions that leave spatial localization to the mask. Experiments with QwenImage-2511 and FLUX.2-dev demonstrate improved editing quality, regional control, and background consistency.
Contributions
- 1
Precise edits, preserved context. A unified training and inference framework combines mask-guided probability paths, a mask-aware objective, and Soft-Poisson De-seaming to improve localization, background consistency, and boundary continuity.
- 2
Masks specify where; prompts specify what. MaskEdit-Benchmark covers general scenes and infographics with region masks and concise instructions that omit explicit object identification and position descriptions.
- 3
Consistent improvements across backbones. Qualitative and quantitative comparisons on QwenImage-2511 and FLUX.2-dev demonstrate stronger regional editing across both image domains.