This study presents a clear and reliable framework for classifying the intent behind scientific citations. It combines multi-model reasoning with concepts from social choice theory. Instead of using a single model, this framework employs three open Large Language Models Gemma, LLaMA, and Mistral. Additionally, we combine their ranked outputs using an exponentially weighted Borda method. By doing so, this approach increases agreement among high-confidence predictions, maintains ranking information, and produces stable, high-quality supervision signals. Consequently, it boosts reliability while remaining transparent. To create a strong experimental basis, we built a large, balanced dataset from the UnarXive corpus, which contains structured full-text scientific articles and citation networks. First, we automatically pulled citation contexts and organized them within a DuckDB-based analytical setup. Then, we rebalanced the dataset across rhetorical categories to enhance representativeness and minimize bias. Finally, we categorized each citation context into one of five roles: background, methodology, comparison, extension, or critique. As a result, the resulting dataset provides a robust foundation for training and evaluation. We trained a SciBERT classifier using these ensemble-generated annotations and tested it on a five-category citation intent classification task. The model achieved a macro F1-score of 0.83, an outstanding result for this type of classification. Indeed, this level of performance shows strong reliability given how challenging it is to differentiate closely related citation functions. Moreover, it demonstrates that combining multiple models produces valuable and distinct supervision signals, capturing subtle rhetorical and semantic patterns that single models often overlook. Furthermore, the framework enhances interpretability. Specifically, the explicit weighting system clarifies how each model contributes to the final outcome. In addition, the deterministic tie-breaking method ensures the outputs are consistent and reproducible. Taken together, these design choices maintain explainability without sacrificing effectiveness.
M. Barchane, Saad Belefqih, El habib Ben lahmar et al.· Algorithms· 0 citations
Schema matching plays a crucial role in data integration by aligning attributes from heterogeneous data sources. However, traditional approaches often fail to capture semantic relationships between attributes with different naming conventions. To address this limitation, we propose SchemaML, a supervised machine learning approach that combines semantic word embeddings with a Random Forest classifier. The proposed method follows a structured pipeline including column extraction, embedding generation using pre-trained models, pairwise feature construction, and supervised classification. Experiments are conducted on datasets derived from schema.org, covering multiple schema sizes (10, 20, and 50 columns) to evaluate scalability and robustness. SchemaML achieves an accuracy of up to 94%, with precision and recall exceeding 93%, outperforming traditional rule-based and machine learning baselines under identical experimental conditions. Additional evaluations demonstrate robustness to noisy data, including typographical errors and abbreviations, with performance remaining stable under 10% noise injection. These results highlight the effectiveness of combining semantic representations and ensemble learning for scalable and automated schema matching in real-world data integration scenarios.
Mohamed Raoui, Moulay Hafid El Yazidi, A. Zellou· 2026 6th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.