Harness Engineering for Multi-Agent Data Visualization with Small Language Models
Abstract
Turning a natural-language question into a correct, publication-ready data visualization usually takes several rounds of coding, inspection, and debugging. Large language models can draft plotting code, but a single-shot model call has no way to run the code, look at the resulting image, recover from execution errors, or decide whether the chart actually answers the user. We demonstrate a visualization harness: an engineered scaffold that wraps a team of seven specialized agents in a bounded Plan–Execute–Verify–Refine control loop. The harness, not any single model, supplies the capabilities that make the system reliable—a sandboxed execution environment with timeouts, vision-based verification of the rendered PNG, error-driven routing between re-planning, re-coding and debugging, per-agent model assignment, and persistent run state for resume and reproducibility. Because the harness owns reliability, the underlying models can be small: a hybrid configuration that runs four agents on local small language models (3–8B) and routes only the three failure-prone agents to inexpensive cloud models reaches a 100% verification pass rate across five heterogeneous datasets at roughly one-third the cost of an all-cloud baseline. The live demo, built on a lightweight FastAPI web UI, lets attendees upload a dataset, type a query, and watch agent steps and intermediate charts stream in real time over Server-Sent Events until a verified visualization appears.