Oct 2026· Proceedings of the 14th Nordic Conference on Human-Computer Interaction· 0 citations· 21 references
TL;DR
BranchVis is presented, a prototype that represents prompts and generated images as nodes in a tree-based visualization, enabling direct navigation, unambiguous revisitation, integrated comparison, and management of complex histories.
Abstract
Generative image editing systems allow users to iteratively modify images through natural language prompts, yet most interfaces present this process as a linear conversational history. This obscures the branching nature of exploratory workflows, making it difficult for users to understand how images evolve, revisit intermediate states, and compare alternatives. In this work, we investigate how interactive visualizations of editing histories can address these limitations. We present BranchVis, a prototype that represents prompts and generated images as nodes in a tree-based visualization, enabling direct navigation, unambiguous revisitation, integrated comparison, and management of complex histories. A within-subject study (N = 12) comparing BranchVis to a conversational interface shows significantly higher usability, lower cognitive workload, and improved support for exploration, navigation, comparison, control, and understandability. These findings highlight editing history visualization as a promising direction for human-AI interaction.
Interactive visual interfaces have become an important means of controlling generative image models, enabling users to manipulate generation through prompts, direct manipulation, and a range of interactions. However, existing techniques are typically presented as independent systems, making it difficult to understand h...
Susie S. Y. Li, Mingwei Li, Remco Chang· 0 citations
Authoring and refining presentation slides is time-consuming in academic and professional settings. Although generative AI lowers the barrier to creating initial drafts, its black-box, one-way workflow often limits fine-grained control. A formative study with 10 frequent presentation authors identified trial-and-error...
Yu Fu, Yong-Qi Kang, Yu-Jia Zhou et al.· Proceedings of the 28th Inte...· 0 citations
A striking pattern emerges: across all methods and modalities, the AI operates strictly as a spatial command executor, with collaborative and model-initiated categories entirely vacant.
Recent generative video editing models enable video content modification (e.g., changing a character) but target short clips. Extending them to full multi-shot videos requires tedious work to locate relevant content across shots, segment it into clips, craft context-aware editing prompts for each clip, and repeatedly a...
Boyu Li, Yu-Qian Zhou, Duo-Tun Wang et al.· 0 citations
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to...
Cheng Yang, Chu-Fan Shi, Hui-Juan Wang et al.· 0 citations
Compositing multiple visualizations into a coherent whole remains challenging due to the vast design space and the need to balance the coverage of task-relevant data insights (e.g., trends and outliers), perceptual clarity, and aesthetic quality. In this paper, we present VisPuzzle, a task-aware method that formulates...
Zheng Wang, Zhiyang Shen, Lingyun Yu et al.· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.