It is found that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated references match real papers in bibliometric databases.
Abstract
The spread of scientific knowledge depends on how researchers discover and cite prior work. Large language models (LLMs) now add a new layer to this process, but their alignment with human citation practices across domains remains unclear. Here, we compare human citations with GPT-4ogenerated reference suggestions produced from paper metadata and abstracts. Analyzing 274, 951 generated references for 10, 000 focal papers, we find that LLMs systematically reinforce the Matthew effect by favoring highly cited papers, with field-specific variation in the rate at which generated references match real papers in bibliometric databases. Generated references diverge from groundtruth reference lists by favoring more recent papers, shorter titles, and smaller author teams. Yet they remain semantically aligned with focal-paper content at levels comparable to human references, reproduce similar local citation-network structure, and reduce author self-citations. These results show that LLMs can generate content-relevant bibliographic suggestions from parametric knowledge alone, but that they also amplify dominant citation patterns. As such tools become routine in research workflows, they may reshape how scientific communities discover, prioritize, and build on prior work.
Citation function classification plays a crucial role in understanding the relationships between scientific publications and advancing bibliometric analysis. This study presents one of the first comprehensive evaluations of multiple state-of-the-art (SOTA) large language models (LLMs) for citation function classification, achieving new SOTA results on the ACL-ARC dataset. We systematically compare five models (Mistral 7B, Orca 2-7B, LLaMA 3.1-8B, Falcon 7B, and SciBERT) across zero-shot, few-shot, and fine-tuning approaches. Our fine-tuned Falcon 7B model achieves a 73.3% macro F1 score on ACL-ARC, representing a significant improvement over previous methods. Additionally, we introduce AC3, a novel dataset featuring a seven-category annotation scheme that distinguishes between neutral acknowledgments and explicit evaluative stances (more opinion-oriented citations - criticizing, complimenting, contradicting). The dataset is implemented across four context extraction variants to systematically evaluate the impact of contextual scope on classification performance. We also provide detailed analysis of model performance, experimental configurations, and limitations to guide future research in this domain. To our knowledge, this is one of the first studies dedicated to comprehensive model comparison for citation function classification, addressing a gap identified in recent surveys.
Daniel Vodička, Jakub Šmíd, Pavel Král et al.· International Conference on...· 0 citations
Large language models increasingly mediate access to information, and a growing body of work asks whether they cite the sources behind their outputs. That work treats citation as a verification device and applies it to textual documents. Scholarly citation serves two further functions, credit and provenance, and it applies to data as much as to text. This paper argues that data citation for large language models is an open challenge, distinct from document-level citation grounding and harder to solve. We ask how such models should cite data so that outputs stay verifiable, provenance stays traceable, and credit reaches data creators and curators. We set out three research directions. Training data attribution has to turn influence estimates into references for corpora absorbed into model parameters. Data citation at inference time has to identify datasets, subsets, and query results at the right granularity and fixity. Citing knowledge graph facts has to define what a reference to a single triple denotes and how credit propagates along provenance. Progress on all three depends on joint work across the database, information retrieval, knowledge representation, and artificial intelligence communities.
Understanding the flow and evolution of scientific knowledge is essential for assessing research impact. Existing citation analysis methods mainly focus on citing authors'subjective intents, failing to consistently characterize cited papers'knowledge contributions. This study proposes the Knowledge Contribution Taxonomy (KCT), derived from the Scientific Research Logic Model, which identifies the type of knowledge a cited paper contributes based on the citation context. KCT classifies citations into Method, Resource Tool, Empirical Finding, and Background, further distinguishing core from non-core contributions. We propose a Dual-Path Fusion model for the classification task, which achieves an accuracy of 85.5%, outperforming mainstream large language models. An analysis of 802,202 citations from the ACL Anthology reveals that core knowledge contributions account for only 39.09% of all citations. The core knowledge contribution citation count achieves higher hit rates for award-winning papers than the traditional citation count at all ranking cutoffs, reflecting the value of differentiating knowledge contributions for research evaluation and impact prediction. In dissemination prediction experiments, KCT outperforms citation intent classification, demonstrating its stronger predictive validity for scholarly dissemination. By focusing on the knowledge contributions of cited papers, the KCT can support differentiated research evaluation.
Zhibang Quan, Zhentao Liang, Ming Ma et al.· 0 citations
Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.
Yixuan Liu, Lin Chen, Zhuoqi Liu et al.· 0 citations
Background. Predicting the future influence of scientific papers remains a longstanding challenge in bibliometrics and information retrieval. Traditional regression and embedding-based methods estimate citation counts from textual features but fail to capture the inherently relational and temporal nature of scholarly impact. Objective. This paper proposes a preference-aligned framework for forecasting and generating scholarly influence from paper abstracts and temporal cues. Rather than predicting absolute citation values, we model the relative likelihood that one paper will accrue more citations than another within a shared temporal time frame. Methodology. Impact-DPO integrates temporally informed prompting with direct preference optimization, enabling LLMs to learn comparative influence patterns without explicit graph message passing. We formalize citation forecasting as pairwise preference learning on temporal text-attributed graphs, using publication year as a minimal temporal signal. Experiments were conducted on two large-scale domains, i.e., Computer Science and Physics, spanning three temporal splits (2018–2020). Results. We show that Impact-DPO achieves the highest pairwise accuracy across all splits, with an average pairwise accuracy of 83% (up to \({\approx}88\%\) on individual splits), outperforming SPECTER2+SVR by 12 pp and zero-shot prompting by more than 15 pp. Preference alignment yields an average improvement of roughly +30 pp over binary-classification baselines overall; importantly, supplementary same-backbone Qwen 2.5–7B BC runs remain near chance (about 50–58% across splits), showing that the gain is not explained by backbone capacity alone. A lower regularization parameter ( \(\beta=0.1\) ) consistently produces optimal results, consistent with a low-gap preference regime in citation data. Generative evaluations further reveal that model-generated text aligns more closely with highly cited papers (KS=0.0744, p=0.0065), demonstrating emergent generative alignment with influential scholarly language. Code and data availability. All code, data-preprocessing scripts, and evaluation notebooks are available at https://github.com/parhamouni/impact-dpo.
Parham Hamouni, Ebrahim Bagheri· ACM Transactions on Intellig...· 0 citations
Scientific discovery depends on finding prior literature that shapes what comes next. Existing retrieval systems optimize for relevance and popularity, often favoring central papers over less familiar works that later prove generative. We introduce \textbf{MUSES}, a million-instance benchmark for prospective intellectual-roots retrieval over a fixed 2.33M-paper corpus, with roughly 140K test instances per familiarity tier. To our knowledge, it is the first prospective benchmark at this scale with a shared retrieval task and author-confirmed paper-level root labels. Alongside it, \textbf{CiteRoots} pairs a scalable rhetorical layer over local citation text (LLM judge $\kappa = 0.896$ versus human gold) with a paper-level author-endorsed layer ($n = 1{,}518$ generative-inspiration pairs from 753 focal papers). MUSES organizes difficulty along two axes: a \emph{familiarity} axis spanning CiteNext, CiteNew, and CiteNew-Isolated, and a \emph{functional} axis spanning broad citations, rhetorical roots, and author-endorsed roots. Across 9 method classes, a lean multi-centroid retriever built on SPECTER2 is strongest. Hit@100 falls from 0.534 on CiteNext to 0.424 on CiteNew, 0.205 on rhetorical CiteNew, and 0.171 on author-endorsed CiteNew, a $3.1\times$ decline. In a registered eight-lens full-test audit, roughly half of broad-tier test instances remain unsolved at K=1{,}000. Rhetorical role and author endorsement are distinct: the same judge agrees with endorsement at $\kappa = 0.037$. We release MUSES, both CiteRoots layers, and a distilled open companion judge for future work on prospective retrieval and intellectual roots.
Rohan Pandey, Sunjae Kwon, Hong Yu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.