R-Select, a robust and scalable framework that optimizes data selection with 30 distinct quality metrics, introduces a novel hierarchical optimization strategy that consistently outperforms both heuristic baselines and model-based methods, offering a robust solution for high-quality data curation.
Xin Gao, Xiao-Yang Wang, Yun Zhu et al.· Proceedings of the 32nd ACM...· 0 citations
Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited emphasis on mechanism understanding, evidence-grounded reasoning, and hypothesis evaluation. To address this, we introduce SPARK (Scientific Paper Abstracted Reasoning sKeleton), a paper-oriented synthesis framework built on Sci-Base, a large-scale corpus of research papers spanning 10 scientific disciplines. Instead of directly converting papers into question-answer pairs, SPARK treats the claim-evidence-derivation structure of a paper as the fundamental unit of reasoning synthesis. Specifically, SPARK (1) distills each paper into a compact reasoning skeleton capturing its central claims and supporting evidence, enabling self-contained question generation, and (2) synthesizes reasoning tasks from four scientific perspectives: mechanistic reasoning, hypothesis falsification, quantitative derivation, and boundary calibration. A final consistency verification stage further removes unsupported or contradictory outputs. Using this framework, we construct Spark-234K, a scientific reasoning dataset with substantially higher difficulty and diversity than existing resources. Experiments show that Spark-234K consistently outperforms existing scientific reasoning datasets while achieving stronger performance with significantly fewer training samples.
The transition from architecture-centric scaling to data-centric refinement has established high-quality data as a critical determinant of Large Language Model performance, particularly for complex reasoning and instruction following. However, effective data selection remains a persistent bottleneck: simple heuristic filters often fail to capture multifaceted data features (e.g., reasoning depth and diversity), while advanced model-based scoring methods typically prioritize isolated quality dimensions, failing to provide a holistic assessment. To address these challenges, we propose R-Select, a robust and scalable framework that optimizes data selection with 30 distinct quality metrics. Recognizing that optimizing such a high-dimensional feature space is non-trivial, R-Select introduces a novel hierarchical optimization strategy. This approach structurally decomposes the search problem by first clustering correlated metrics into functional groups based on statistical dependencies. It then executes a two-stage optimization process: performing intra-group refinement to maximize local representational power, followed by inter-group integration to balance global quality domains. Crucially, to ensure computational efficiency, R-Select employs a low-resource proxy strategy, utilizing a lightweight model on a small data subset to learn an optimal selection policy that is transferable to the target model. Extensive experiments demonstrate that R-Select consistently outperforms both heuristic baselines and model-based methods, offering a robust solution for high-quality data curation.
Xin Gao, Xiaoyang Wang, Yun Zhu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.