As image generation and editing models become increasingly capable, a central challenge is ensuring that their outputs faithfully reflect user intent while making their internal behavior understandable and amenable to precise intervention. This work develops methods to improve the interpretability and controllability of these models, with a focus on compositionality, knowledge localization, and fine-grained controllability.
For compositional generation, we first identify limitations in CLIP-based text conditioning that impair attribute–object binding and introduce lightweight projection-based interventions that improve compositional fidelity while preserving image quality. We then develop an agentic framework that constructs compositionally contrasted training examples and uses distance-aware preference optimization to improve complex prompt following.
We next study how semantic knowledge is represented within diffusion transformers. By tracing textual information across transformer blocks, we localize components responsible for specific visual concepts and validate their causal roles through targeted interventions. These findings enable efficient, selective personalization and knowledge unlearning while preserving unrelated capabilities.
Finally, for instruction-based image editing, we develop a framework that disentangles multi-part instructions into continuous controls, allowing users to precisely adjust the strength of each individual edit and iteratively reach their desired output—an essential capability for interactive editing interfaces and fine-grained user control.
Together, these works connect mechanistic understanding, targeted model adaptation, and intuitive user control to advance image generation and editing models that are more faithful, interpretable, efficient, and controllable.
Arman Zarei is a Ph.D. student in the Department of Computer Science at the University of Maryland, College Park, advised by Prof. Soheil Feizi. His research lies at the intersection of computer vision and generative AI, with a particular focus on improving the interpretability and controllability of image generation and editing models, especially in compositionality, knowledge localization, and fine-grained user controllability.

