Large language models (LLMs) are increasingly used through terminal user interfaces (TUIs), yet task completion alone does not capture how difficult an interface is to understand and operate. Existing human assessments and LLM-generated ratings or reports do not provide repeatable measurements of interaction effort gro...
Rui-Da Hu, Yuan-Hao Wang, Chao Peng et al.· 0 citations
This work surveys real-world TUI applications, turns them into a headless benchmark spanning ratatui/Rust, bubbletea/Go, textual/Python, and ink/TypeScript, and compares four frontier LLMs with random exploration, finding no model dominates.
Chao Peng, Ruida Hu, Ajitha Rajan et al.· 0 citations
This work proposes TraceDev, a multi-agent framework for automated software development grounded in use cases that contain multiple functional points and complex semantics, and demonstrates the effectiveness of TraceDev in repository-level code generation from requirements.