RaivenTracks is presented, a workflow-aware extension of the Raiven DSL-mediated visualization pipeline that treats validated visualization specifications as persistent, branchable checkpoints and frames branchable conversational visualization history as a step toward provenance support for future scientist-in-the-loop oversight of AI-driven scientific workflows.
Abstract
As AI agents increasingly participate in scientific workflows, scientists are shifting from direct authorship toward oversight, inspection, and steering. LLM-driven visualization systems are a promising interface for this hand-off, yet they remain largely stateless, forcing users to reconstruct context across refinements and offering little support for revisiting prior decisions or exploring alternatives. We present RaivenTracks, a workflow-aware extension of the Raiven DSL-mediated visualization pipeline that treats validated visualization specifications as persistent, branchable checkpoints. Because each checkpoint is a verifiable RaivenDSL specification rather than a dialogue transcript, restoring a node recompiles a known artifact rather than re-interpreting prior context. RaivenTracks contributes a two-level state management architecture that pairs a persistent, branchable version tree with a fine-grained undo/redo stack over runtime visualization settings, across both InfoVis and SciVis backends. A formative pilot study with three visualization researchers shows early promise, with all participants adopting the version tree for branching and recovery, and surfaces design directions for tree navigation, node labeling, and scalability that inform a planned controlled comparison against Raiven without version history. We frame branchable conversational visualization history as a step toward provenance support for future scientist-in-the-loop oversight of AI-driven scientific workflows.
The design rests on one claim: most of the credibility of machine-made research can be moved from asking the model to behave to making the non-compliant state unrepresentable, and the system description is a system description written under one rule.
AdaLens is presented, an interactive system for monitoring and steering ongoing runs that combines a storyline-based representation that unifies analytical plans, execution progress, intermediate findings, and data-column involvement with steering interactions grounded in these analytical elements for directional guida...
Yangtian Liu, Yan Miao, Shuhan Liu et al.· 0 citations
Turning a natural-language question into a correct, publication-ready data visualization usually takes several rounds of coding, inspection, and debugging. Large language models can draft plotting code, but a single-shot model call has no way to run the code, look at the resulting image, recover from execution errors,...
MUSE is presented, an interactive meta-agent that enhances user understanding and control of agentic data science systems by dynamically restructuring low-level execution traces into multiple semantic levels that support navigation from high-level overviews to low-level implementation details.
Wei-Hao Chen, Weixi Tong, Yuan Tian et al.· 0 citations
ReFigBench, a benchmark and evaluation framework built on 1,000 real overview figures retrieved from arXiv papers with full provenance, is introduced, exposing the tension between fidelity and editability as the central challenge for practical multimodal document agents.
Li-Yang Fan, Chi Wei, Yi-Tai Li et al.· 0 citations
This work introduces ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations, and describes the release's verification gaps and the pilot's uneven coverage, exclusions, and grading provenance.
David Garg, Ritobrata Sarkar, Ehsan Azarnasab et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.