Research · ICMR 2026
InterFold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples
International Conference on Multimedia Retrieval 2026
InterFold introduces a smarter, more precise way to edit specific details in AI-generated images.
Abstract
We propose InterFold, a framework for learning and applying interpretable semantic manifolds in latent diffusion models, without requiring binary or paired supervision. Existing methods for semantic editing either rely on limited paired data or uncover only coarse, unsupervised directions that fail to capture user-specific, fine-grained attributes. InterFold addresses these limitations by learning a target attribute manifold in the H-space of diffusion models using only a set of positive, unlabeled examples. To edit a new image, InterFold projects its H-space representation toward this learned manifold through test-time optimization, enabling precise, identity-preserving modifications of complex, non-binary concepts. To make these edits effective in modern latent diffusion models, we introduce the Manifold Adapter, a lightweight cross-attention module that transfers semantic intent from edited H-space codes into the generative latent space, without altering the pretrained model. Extensive experiments demonstrate that InterFold achieves superior edit accuracy and identity consistency compared to existing methods, offering a flexible and interpretable solution for high-fidelity semantic image editing.
Cite
@inproceedings{lewi2026interfold,
author = {Lewi, Alexander Vincent and Tan, Rainer and He, Shengfeng},
title = {InterFold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples},
booktitle = {Proceedings of the 2026 International Conference on Multimedia Retrieval},
series = {ICMR '26},
pages = {1861--1869},
year = {2026},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
doi = {10.1145/3805622.3810585},
isbn = {9798400726170},
url = {https://doi.org/10.1145/3805622.3810585}
}