Large language models (LLMs) are increasingly used through terminal user interfaces (TUIs), yet task completion alone does not capture how difficult an interface is to understand and operate. Existing human assessments and LLM-generated ratings or reports do not provide repeatable measurements of interaction effort gro...
Rui-Da Hu, Yuan-Hao Wang, Chao Peng et al.· 0 citations
An enhanced approach based on ITeM is proposed, referred as ITeM-HM, which incorporates specific OpenHarmony system features and achieves a 214% success-rate relative improvement over the original ITeM, which is primarily hindered by OpenHarmony-specific characteristics.
This work introduces Model Automated Deployment Engine (MADE), a dual-agent coordination system that iteratively constructs and validates the deployment artifacts, updates its deployment belief based on execution feedback, and revisits invalid upstream artifacts until the model is successfully served as a ready-to-call...
Yicheng Liu, Bolin Zhang, Weiran Liu et al.· 0 citations
This work proposes TraceDev, a multi-agent framework for automated software development grounded in use cases that contain multiple functional points and complex semantics, and demonstrates the effectiveness of TraceDev in repository-level code generation from requirements.