Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 1262-1273· 0 citations· 18 references
Abstract
The transition from architecture-centric scaling to data-centric refinement has established high-quality data as a critical determinant of Large Language Model performance, particularly for complex reasoning and instruction following. However, effective data selection remains a persistent bottleneck: simple heuristic filters often fail to capture multifaceted data features (e.g., reasoning depth and diversity), while advanced model-based scoring methods typically prioritize isolated quality dimensions, failing to provide a holistic assessment. To address these challenges, we propose R-Select, a robust and scalable framework that optimizes data selection with 30 distinct quality metrics. Recognizing that optimizing such a high-dimensional feature space is non-trivial, R-Select introduces a novel hierarchical optimization strategy. This approach structurally decomposes the search problem by first clustering correlated metrics into functional groups based on statistical dependencies. It then executes a two-stage optimization process: performing intra-group refinement to maximize local representational power, followed by inter-group integration to balance global quality domains. Crucially, to ensure computational efficiency, R-Select employs a low-resource proxy strategy, utilizing a lightweight model on a small data subset to learn an optimal selection policy that is transferable to the target model. Extensive experiments demonstrate that R-Select consistently outperforms both heuristic baselines and model-based methods, offering a robust solution for high-quality data curation.
PPL-Factory is proposed, a simple and interpretable data selection framework that combines task-aware perplexity-based scores and data budget-aware selection criteria that outperforms other state-of-the-art data selection methods using only $1\%$ of the training set.
Hang Zhang, Warren J. Gross· arXiv.org· 0 citations
As the context window of Large Language Models (LLMs) continues to expand, the data required to effectively train and evaluate these capabilities remains underexplored. With existing research primarily focuses on architectural optimization, there is a need for a systematic, data-centric review. This survey bridges this gap by investigating the data foundations of Long-Context Language Models (LCMs). We begin by examining current data strategies alongside their strengths and limitations, mapping the required data to desired model capabilities. Building on this, we explore how targeted training data designs drive core, often interconnected skills such as retrieval, reasoning, and aggregation. Furthermore, we analyze the evaluation landscape, illustrating how selecting appropriate benchmarks is crucial for probing capability boundaries and guiding effective model selection. Finally, we synthesize actionable guidelines for data construction and outline critical future directions to propel the advancement of long-context language models, including quantifying data quality, establishing scaling laws for length distributions, and developing dynamic evaluation frameworks.
Zechen Sun, Yu-Yang Sun, Zhao-yu Su et al.· Transactions of the Associat...· 0 citations
AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa et al.· International Journal of Dat...· 0 citations
DiverValue-Bench is introduced, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions and it is shown that lightweight preference-based fine-tuning with Low-Rank Adaptation and Direct Preference Optimization substantially improves in-domain value alignment while yielding consistent out-of-domain gains.
Yao Liang, Dongcheng Zhao, Feifei Zhao et al.· 0 citations
Multi-Objective Tool-augmented Symbolic Regression (MOT-SR), a unified framework that integrates external analytical tools to extract structural priors and guide equation generation, while jointly optimizing for accuracy, complexity, and generalization via a multi-objective evaluation module that maintains a dynamic Pareto front is proposed.
Boxiao Wang, Runxian Wang, Kai Li et al.· arXiv.org· 0 citations
This paper introduces a comprehensive multi-criteria evaluation methodology designed to assess the capabilities of these advanced computational architectures in handling complex language tasks, focusing on three foundational dimensions: hierarchical reasoning, self-correction mechanisms, and factual consistency.
S. Yam· Journal of innovative resear...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.