Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop...
Li-Chen Zhu, Yueqian Lin, Yi-Heng Wang et al.· 0 citations
Personal LLM assistants (health companions, elder-care agents, accessibility aides) are judged by what they remember about a person: a medication or an allergy mentioned in passing and needed days later, so an eviction policy must decide what the cache forgets.
Li-Chen Zhu, Yueqian Lin, Yi-Heng Wang et al.· Proceedings of the 4th Inter...· 0 citations
SchemaGUI, a template-based benchmark for controllable GUI generation evaluation, synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas can generate thousands of deterministically annotated tasks in seconds without human labeling.
Jiarui Dong, Yin Cai, Zhouhong Gu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.