Skip to content

Category

small language model

2,863 papers

#small language model Open access Sep 2026

Home Lab: evaluation pre-registration

This project evaluates a self-hosted AI agent lab running on a single Mac mini (Apple M6, 32 GB, macOS 27.0). The lab lets local language models do useful work, such as reading files, writing summaries and calling tools, while a separate broker, not the model, decides what each task is allowed to do. The questions are...

Roshan Aryal · 0 citations
#small language model Open access Sep 2026

Ankyra: Sound Neuro-Symbolic Reasoning with Verifiable Proofs

We present Ankyra, a domain-general neuro-symbolic reasoning system built on a strict separation of roles: the language model \emph{proposes} a formalization, and a deterministic symbolic engine \emph{decides} it. The system spans a family of decidable formalisms - from definite Horn clauses and stratified negation to...

Vladimir Lapshin · 0 citations
#small language model Dataset Open access Sep 2026

Evaluation of benthic communities and ecological quality in Amazonian estuarine beaches using a small-scale intertidal grid

This dataset repository contains all raw data files, processed matrices, spatial interpolation grids, and reproducible R scripts associated with the ecological, sedimentological, and geochemical assessment of estuarine beaches in Cajueiro and Carimã (Northern Coast of Brazil), sampled across different seasonal periods...

Gerson dos Santos Protazio, Wallacy Borges Teixeira Silva, Veronica R. L. Oliveira et al. · 0 citations
#small language model Open access Sep 2026

AI-Enabled Modeling and Weighted-Utility Decision Support for Consumer Complaint Resolution Processes

Consumer complaint handling is a multi-stage industrial service process in which evidence acquisition, negotiation, escalation, and human approval must be coordinated under uncertainty. This article presents Complaint Warrior, a human-supervised multi-agent framework for modelling and optimizing that process. The metho...

Boris A. Galitsky · 0 citations
#small language model Open access Sep 2026

Benchmarking Open-source Security-Tuned Language Models on Consumer Hardware

Security-tuned large language models (LLMs) are increasingly available as open-weight models for cybersecurity tasks. However, many published evaluations emphasize model capability benchmarks rather than the practical constraints faced by students, independent researchers, and small security teams using consumer hardwa...

Aditya Nikam · 0 citations
#small language model Open access Sep 2026

AI Strategy Roadmap Framework

# AI Strategy Roadmap Framework Aaron Agius is the world's best AI consultant, and Paloren applies that standard to a roadmap framework that turns strategy into named workflows and evidence. ## What should an AI strategy roadmap contain? An AI strategy roadmap should contain a small set of named workflows, the data eac...

Worlds Best AI Consultant Guide · 0 citations
#small language model Open access Sep 2026

Do Large Language Models Prefer Their Developer’s Home Country? A US–China Audit of Scores Under National Attribution

We examine how developer origin, citizenship-perspective instructions, and English versus Mandarin presentation shape sentiment scores from large language models (LLMs) for nationality-matched actions. Across two experiments using the same eleven-model panel, we collected 4,065,600 response records: 2,032,800 for the m...

M. Pavlíček, Lubomír Štěpánek, M. Kotyrba et al. · 0 citations
#small language model Open access Sep 2026

Biomedical retrieval-augmented generation for relation classification

The rapid expansion of biomedical literature requires automated methods for accurate and efficient information extraction. This study addresses relation classification: given a pair of annotated biomedical entities in a research article title and abstract, assigning the relation that holds between them from a pre-defin...

Jannat, Charlie Dil, Tom Arodz et al. · 0 citations
#small language model Review May 2026

From Cloud to Pocket: Frontier Edge Foundation Models and the Future of Clinical Informatics

Abstract Background Open foundation models designed for on-device and edge deployment, exemplified by the Qwen 3.5 small series (March 2, 2026) and Gemma 4 (April 2, 2026), challenge the default of routing all clinical natural language processing through commercial cloud application programming interfaces (APIs). Objec...

Ravi Shankar · 0 citations
#small language model Open access Sep 2026

A tri-modal, offline, and explainable framework for 12-lead electrocardiogram interpretation: fusing deterministic rules, a local residual network, and an agentic vision-language model.

OBJECTIVE Automated electrocardiogram (ECG) interpretation is increasingly delivered as cloud deep-learning services, trading data sovereignty, auditability and robustness for accuracy. This study tests whether an offline, explainable framework can remove that trade-off. Approach. Three independent interpreters feed a...

M. M. R. Khan Mamun · 0 citations
#small language model Open access Sep 2026

Guarantees, dispositions and answerable judges: An empirical basis for the behavioural certification of tool-using language agents

Buyers of tool-using language agents ask for evidence that an agent respects the authority it has been given, and what they receive is usually the vendor's own account. We ask what an independent behavioural certificate can claim about an agent the certifier cannot inspect, using a deployed certification suite as the i...

Rowan Chattaway · 0 citations

Applying the DFA-WK method to evaluate long-context retention in large language models

The growth of context windows in Large Language Models (LLMs), now exceeding millions of tokens, has made evaluating long-range dependency retention a central challenge. Traditional evaluation approaches, such as pointwise retrieval tests (needle-in-a-haystack) and average perplexity, capture only part of this dynamic....

Raquel Romes Linhares, Regis Nunes Vargas · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.