Oct 2026· Proceedings of the 28th International Conference On Multimodal Interaction· 0 citations· 19 references
Abstract
Authoring and refining presentation slides is time-consuming in academic and professional settings. Although generative AI lowers the barrier to creating initial drafts, its black-box, one-way workflow often limits fine-grained control. A formative study with 10 frequent presentation authors identified trial-and-error anxiety, invisible intermediate decisions, cross-slide inconsistency, ambiguous spatial references, and fear of irreversible edits. We present ECHO, an interactive slide-refinement system that combines natural-language instructions with direct visual selection. ECHO translates multimodal intent into schema-constrained operation plans through a Plan-Confirm-Execute loop, maintains global style and user habit memory, routes spatially ambiguous requests to a vision-language model, and supports byte-exact rollback. We further introduce CoEdit-Eval, a multi-level framework for evaluating intent mapping, spatial grounding, execution safety, and rendered visual quality. Across four foundation models, ECHO raises Target Hit@1 from 0% for text-only baselines to 55–85%. A within-subjects study with 14 participants shows a 25.8% reduction in completion time and a 20.8% reduction in NASA-TLX workload. These results demonstrate how explicit execution boundaries can make AI-assisted document refinement more controllable, transparent, and reversible.
ReDeck is proposed, a step-level render-grounded refinement framework that decomposes slide revision into atomic edit actions and returns renderer-derived observations after each step, turning refinement into"one edit, one observation."
Mu-Zhao Tian, Ze-Zi Zeng, Yi-Fan Yang et al.· 0 citations
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield dr...
Aman Tyagi, Hemanth Boinpally, Jonathan Chen et al.· 1 citation
This work introduces PROS (Proactive Refinement Of Scientific Posters), which separates epistemic initiative from behavioral authority: the system can surface source-grounded candidate problems, while users decide which become repair goals and whether resulting changes are committed.
Xingda Lyu, Hong-Ling Lu, Xin-ye Luo et al.· 0 citations
Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only \emph{flat}, absolutely positioned elements, so agents must recompute coordinates and routinely break layouts; and design has no un...
Jooyoung Jang, Taegyeong Lee, Jihyeon Park et al.· 0 citations
BranchVis is presented, a prototype that represents prompts and generated images as nodes in a tree-based visualization, enabling direct navigation, unambiguous revisitation, integrated comparison, and management of complex histories.
Rifat Mehreen Amin, Alexander Nuss, M. Dutoit et al.· Proceedings of the 14th Nord...· 0 citations
Tool-augmented vision-language models increasingly"think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has shown that returned images contribute little, raising the question: if pixels do not car...
Jiahao Shao, Yuanbo Yang, Yi-Yi Liao et al.· 3 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.