Back to #small language model
#small language model Open access

What Do Little Machines in Spring 2026 Know About Buddhist History?

Aug 2026 · Yin-Cheng Journal of Contemporary Buddhism · 0 citations · 14 references

TL;DR

This paper test representative models regarding their knowledge of Chinese and Japanese Buddhist history on consumer hardware and finds that, in Spring 2026, the Qwen series emerges as a winner for models in the 30B range, while the larger Kimi2.5 models lead in a cloud-based setup.

Abstract

The availability of open-weights language models has led to a profusion of language models trained for particular tasks. These relatively small, but increasingly powerful, models run on consumer hardware and can be used freely for the price of electricity. Because the training data for most models are not made public, it is difficult to say how much knowledge about a given domain is embedded in any model; even if we knew more about the training data, it is not always clear how much of it a model can externalize. This paper is a first step towards developing a benchmark to test the historical knowledge of language models in the domain of Buddhist Studies. We test representative models regarding their knowledge of Chinese and Japanese Buddhist history on consumer hardware. A test bank, subdivided by historical period, with multiple choice questions (MCQs) is used with different model series. Ranking the answers produces a “winner” that “knows” most about Buddhist history. We find that, in Spring 2026, the Qwen series emerges as a winner for models in the 30B range, while the larger Kimi2.5 models lead in a cloud-based setup.

Read PDF

Similar papers

Preprint Jul 2026

Can a Language Model Learn Facts Continually in Its Weights?

This work follows invented facts written into Qwen3 models from creation through sequences of twenty to one hundred later writes, using held-out questions of five types, and finds that facts can be behaviourally forgotten without being erased.

Charles O'Neill · 0 citations
Open access Aug 2026

Can Knowledge Be Translated (by a Machine)?

This paper concludes that machines may assist translators but that they will not, by principle, be able to reach an almost perfect level and object to the huge amount of money spent on software development for systems with that objective.

H. Götzsche · 0 citations
Review Jul 2026

LLM for the development of FCM

This article is about the development of a fuzzy cognitive map using a local large language model, and the model is thoroughly tested; Qwen2.5-32B is used and the data is extracted from hotel reviews from TripAdvisor and a fuzzy cognitive map is trained and evaluated.

Alexis Kafantaris · 0 citations
Open access Jul 2026

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

Abstract Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabilities, mechanisms, and impacts. Yet these debates remain structured by persistent folk theories—intuitive, informal explanatory models that guide attitudes and actions. Deflationary slogans (“just autocomplete,” “stochastic parrots,” and “average of the internet”) and anthropomorphic framings (“emergent agents” and “proto-minds”) each capture genuine features of current systems but mistake those features for the whole. This Perspective proposes a minimal working model of LLM-based systems centered on four distinctions: between pretraining and deployed systems; between the learned distribution and particular samples; among parametric, contextual, and external memory; and between task competence and agency. The model is used to diagnose six misconceptions about LLMs: next-token prediction, regression to the mean, training-data regurgitation, model memory, alignment, and understanding. For each, the analysis identifies what the misconception gets right, which distinctions it conflates, and what follows for capability evaluation, system design, and governance. Applied to publisher AI policies as governance case studies, the framework shows both how policy language can conflate these distinctions and how such errors can be corrected. The model thereby avoids the parrot–mind binary by treating LLMs as simulators of discourse and task performance, offering a diagnostic toolkit for locating and correcting the errors these folk theories perpetuate.

Zhicheng Lin · 0 citations
Preprint Jul 2026

Probabilistic"Copies"in Generative AI Models

It is argued that copyright law will likely take a functional approach to the question, finding that LLMs contain a copy of a particular work only if it is straightforward to extract that work in outputs.

Mark A. Lemley, A. F. Cooper · 0 citations
Preprint Jul 2026

A dataset of rated conceptual arguments

Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we call conceptual questions: questions for which no ground truth is realistically accessible and no widely accepted resolution methodology exists, but on which progress can still be made by debating arguments. Most philosophical questions are of this kind, as are central components of questions in AI safety, decision theory, and social choice. Our approach is based on the view that while bottom-line conclusions on such questions are hard to evaluate, individual contextualized arguments can be evaluated far more reliably. We therefore introduce a dataset of 951 argumentative critiques of 442 position texts, spanning topics from AI safety and decision theory to ethics and politics, with 1,458 ratings by six expert raters along dimensions including centrality, strength, correctness, and clarity. We propose two scoring functions and benchmark a range of models. Performance tracks general capability rankings.

Emery Cooper, Caspar Oesterheld, Linh Nguyen et al. · 0 citations

Related blog posts