Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.
Yong Peng, Qing-Shui Gu, Li-Ya Zhu et al.· 0 citations
ZenGen, an integrated framework for measuring, internalizing, and grounding social intelligence, and Actio, a harness-controlled inference architecture that routes four typed supports into reasoning demonstrate the effectiveness of typed runtime support.
ZenGen Team, Xiang Ao, Jingping Bi et al.· 0 citations
This work presents Meta's network architecture and software stack designed to support one of the world's largest RoCE fabrics, currently connecting over 100,000 GPUs across multiple datacenter buildings, and introduces a scalable initialization strategy that reduces startup times by 11× via eager process group creation and O(N) topology discovery.
Hongyi Zeng, Min Si, Pavan Balaji et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.