Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading workflows. We introduce SciDocBench, a workflow-centered benchmark for scientific document understanding. It contains 124 expert-authored and difficulty-screened questions organized into seven research-assistant capability groups and 19 subtasks across five scientific domains. Each question is instantiated under four matched conditions combining English or Chinese questions with all-images-first or interleaved document representations, yielding 496 evaluation instances for controlled analysis. The strongest evaluated system achieves only 62.6/100, with pronounced weaknesses in document perception, evidence grounding, verification, and cross-document reasoning. To translate these diagnostics into scalable training signals, we introduce SciDocIR, a typed evidence-graph representation that preserves scientific document objects, layout and cross-reference relations, and provenance. Building on SciDocIR, we construct SciDocDataset, comprising approximately 15K supervised fine-tuning samples and 8K reinforcement-learning samples across 14 verifiable subtasks. Together, SciDocBench, SciDocIR, and SciDocDataset form an evaluation-to-training framework for diagnosing and improving scientific-document assistants. The project page is available at https://github.com/InternLM/SciDocBench.
Shenxi Wu, Yu-Hong Liu, Haosong Zhang et al.· 0 citations
Tabular Anomaly Detection (TAD) plays a fundamental role in securing real-world applications. Despite rapid advances in TAD, the prohibitive cost of human-centric label annotation remains a primary bottleneck for large-scale production systems. To alleviate this bottleneck, we propose a novel ''coarse-to-fine'' label annotation pipeline to improve labor efficiency through a coarse-grained label annotation and fine-grained human verification. Specifically, Large Language Models (LLMs), with their strong cross-domain capabilities, serve as a promising solution for the coarse-grained annotation stage. However, effectively generalizing LLMs to coarse-grained annotation remains challenging due to the inability to ground semantic priors in rigorous deduction, as well as the overfitting risks inherent in single-domain fine-tuning. Accordingly, we introduce TaDGeneral, a large-scale cross-domain corpus constructed by fusing deductive reasoning paths from diverse domains. This design bridges the reasoning gap while preventing the memorization of local shortcuts. Building upon this, we develop TaDFM, a foundation model tailored to internalize generalizable deductive logic for effective zero-shot annotation. Extensive experiments on both public and large-scale real-world TAD datasets demonstrate the superiority of TaDFM over representative methods, with its practical value further validated by an industrial case study. Code: https://github.com/cshhzhao/TaDFM.
Haihong Zhao, Aochuan Chen, Miao Peng et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.