Skip to content
Review Open access

Piercing a Methodological Bubble: Large Language Models and the Social Sciences

Jul 2026 · Fudan Journal of the Humanities and Social Sciences · 1 citation · 24 references

TL;DR

Recognizing this clarifies where LLMs are most useful: as tools for abductive reasoning and theory development, not autonomous instruments of measurement or inference in social science research.

Abstract

Large language models (LLMs) are increasingly used for text classification, survey simulation, and causal analysis in the social sciences. Yet many applications make a category error by treating systems trained to generate plausible language as if they directly measure social reality. This article argues that the problem is conditional rather than categorical. Autoregressive LLM outputs can have a definite relationship to social phenomena only when structured human involvement grounds, verifies, and anchors them to observable evidence. Without such grounding, they remain probabilistically plausible continuations shaped by linguistic patterns rather than by the political world itself. This distinction matters across use cases. Zero-shot generative coding and synthetic response generation are epistemically weakest; treating model outputs as observed variables creates similar problems. Supervised fine-tuned classification is sounder because human-labeled data provides an explicit connection, though the grounding comes from the labeling process rather than from the model alone. Recognizing this clarifies where LLMs are most useful: as tools for abductive reasoning and theory development, not autonomous instruments of measurement or inference in social science research.

Read PDF

Similar papers

Review Jul 2026

Analyzing and Correcting Benevolence Bias in Large Language Models

Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.

Yuanzi Li, Jun-Hao Wang, Minghui Liu et al. · 0 citations
#natural language process... Preprint Sep 2026

From Echo Chambers to Epistemic Monoculture: Large Language Models Present Temporally Contingent Partisan Alignments as Knowledge

Large language models (LLMs) are rapidly becoming an interface between citizens and political information. They are often regarded as"a better Google."While this analogy might work for some instances, it is unintuitively problematic for democratic politics. A search engine retrieves human-authored documents, while a language model generates novel text that necessarily embeds invisible framing decisions. Because conveying knowledge involves framing, a system that generates answers cannot serve as a neutral conduit to"all human knowledge."Instead, these systems are becoming a new kind of political intermediary. Mechanistic evidence shows that partisan identity is encoded as a locatable geometric direction inside the Llama 3.1 8B model, and that alignment training masks rather than removes this structure. Building on that evidence, we present steering experiments that exploit a model's training cutoff in 2024. This cutpoint auspiciously falls just before a dramatic realignment in American politics marked by the second Trump administration and the MAHA transformation of health politics, providing us with a natural experiment. We find that the model presents temporally contingent partisan alignments as knowledge, with no mechanism for distinguishing fact from opinion. This reality moves the information environment beyond the echo chamber toward an epistemic monoculture where language models, purporting to summarize"all human knowledge"are, in actuality, simply magnifying the cultural and partisan divides inherent in their training data.

Wend K. Tam · 0 citations
Sep 2026

Assembling Topic Models: Material Political Economy and the Genealogy of an Algorithm.

Natural Language Processing (NLP) technologies-ranging from topic models to today's large language models like GPT-have rapidly entered the social sciences, reshaping methodological practice. Yet researchers often overlook the stark political-economic contrasts between academia and the AI research-industry symbiosis. Identical algorithms, once embedded in different institutional settings, acquire different meanings and standards of evaluation. This paper shows the divergence by examining topic modeling, a classical NLP technique in computational social science. Social scientists grapple with the instability of applying topic models to the same corpus, whereas in the AI industry such variability matters little, given different evaluative priorities. Through a comparative analysis of topic modeling's trajectory across AI and social science, I show how organizational contexts and goals shape the development of the same algorithms, and why framing instability as a purely technical issue is problematic in the social sciences. The findings reveal that algorithms are not simply technical tools but products of material political-economic regimes. Recognizing this, I argue that STS scholars have a vital role to play in computational social science: not only by critically examining and developing methods, but also by interrogating the material-political-economic regimes in which algorithms are enacted, and by working toward more just alternatives.

Unknown authors · 0 citations
#artificial intelligence Review Sep 2026

Mapping the Emerging Social Science of Large Language Models

Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.

Yi Yang, Xiao Jia, Ze-Yu Dong et al. · 0 citations
Preprint Aug 2026

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.

Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dutch Books for Language Models

People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that these forecasts fall out of a coherent world model. In this paper, we evaluate the coherence of language model probabilistic forecasts through a procedure that builds on a theorem due to de Finetti. We elicit forecasts from language models across events generated from stock returns data. We then use linear programs to compute the largest Dutch-book profit - the profit an arbitrageur could guarantee by betting against model-generated probabilities - which we use as a measure of incoherence. Our procedure does not require outcome labels, so we can evaluate coherence even in settings where outcomes are not observed or have not yet resolved. We find substantial evidence of incoherence in language model forecasts. Such incoherence increases when there are richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude. We conclude by discussing how alternative training strategies may improve probabilistic coherence.

Isaiah Andrews, Suproteem K. Sarkar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.