Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.
Liangcai Su, Zhaopeng Feng, Zhuo Chen et al.· 0 citations
Background Generative artificial intelligence chatbots are increasingly used as sources of consumer health information. Although their performance has been examined in several medical conditions, evidence specific to hand, foot, and mouth disease (HFMD) remains limited. Objective To compare the safety, accuracy, empathy, information quality, and readability of HFMD-related responses generated by five publicly accessible chatbots. Methods In this exploratory cross-sectional study, 20 researcher-developed English-language prompts were constructed from authoritative public-health sources, Google Trends topic mapping, and caregiver-informed wording refinement. ChatGPT-4o, Gemini 2.5 Pro, Copilot, Doubao, and DeepSeek-V3.2-Exp were evaluated between April 2 and April 5, 2026. Five trained reviewers independently assessed responses against predefined reference standards using safety and accuracy criteria, an empathy scale, DISCERN, EQIP, the Global Quality Score, and JAMA benchmarks. Six formula-based readability indices were calculated. Paired comparisons used Friedman and Cochran Q tests, with prespecified post-hoc procedures and Benjamini-Hochberg correction. Results Unsafe-response rates ranged from 5.0 to 15.0%, with no statistically significant difference detected among chatbots (p = 0.797). No statistically significant inter-model differences were detected for accuracy, empathy, DISCERN, EQIP, JAMA, or Global Quality Score in the 20-prompt set; because the study was exploratory and was not powered to establish equivalence, these findings do not demonstrate comparable or interchangeable performance. All six readability indices differed significantly among chatbots (p = 0.030 to <0.001). ChatGPT and Doubao generally produced lower estimated grade-level complexity than Gemini and DeepSeek. Ten potentially unsafe or misleading responses were identified, mainly involving overgeneralization of EV71 vaccine protection, hand-hygiene qualification, and disinfection advice. Conclusion Across a limited set of standardized English-language prompts, the five chatbots often generated coherent HFMD information, but occasional safety-relevant inaccuracies and readability barriers remained. These systems may support general information seeking, but their responses require cautious interpretation and should not replace individualized professional advice.
Zhao-Le Gong, Yan Na, Yi Guo et al.· Frontiers in Public Health· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.