Despite significant progress in general visual question answering and cross-modal understanding, multimodal large language models still face a pronounced gap in evaluation for complex reasoning within the mechanical engineering domain. Existing benchmarks predominantly focus on rudimentary tasks such as drawing recogni...
Tengyue Wang, Kang An, Chenxu Du et al.· 0 citations
benchmark provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning, and provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbol...
Xinqi Yang, Kang An, Tengyue Wang et al.· 0 citations
MMArch is introduced, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers, and error analysis shows that failures concentrate in applying principles and combining evidence across figures rather than in locating it, pointing to substantial he...
Chenxu Du, Kang An, Tengyue Wang et al.· 0 citations
Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoni...
Yuanchi Zhu, Kang An, Tengyue Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.