Skip to content

Category

large language models

2,073 papers

#large language models Dataset Open access Oct 2026

CORE-LLM-Bench: A Controlled Neurosymbolic Benchmark for Ontology-Grounded Reasoning in Large Language Models

CORE-LLM-Bench is a controlled neurosymbolic benchmark for evaluating ontology-grounded reasoning in large language models. Version 1.1.1 is a corrective release that supersedes v1.1.0 for future use. It contains 9,048 unique question-hop instances: 6,032 binary question-answering (BQA) instances and 3,016 open-ended q...

Julie Loesch, Sara Falahatkar, Nicole Kilk et al. · 0 citations
#large language models Open access Oct 2026

Kayenta AI Canary Judge: a model-agnostic remote judge and benchmark for automated canary analysis

A configurable remote judge for Kayenta that produces a canary verdict from a large language model or a vision-language model, three explicit hybrid policies that combine it with the genuine NetflixACAJudge verdict, a tuned statistical ensemble baseline, and a seeded nine-family fault-injection benchmark of 180 labelle...

Vlad-Stefan Dieaconu, Răzvan Rughiniș, Andreea-Loredana Dieaconu-Usurelu et al. · 0 citations
#large language models Dataset Open access Oct 2026

Result data for: Can Large Language Model and Vision-Language Model Judges Enhance Automated Canary Analysis? A Per-Family Evaluation Against Kayenta's Statistical Judge

Every result artefact behind the paper, in the directory layout the analysis code expects, together with the scripts that recompute the paper's tables from them and the analysis code that turns them into every number the article reports. The testbed that produced the measurements is the kayenta-ai-canary-judge reposito...

Vlad-Stefan Dieaconu, Răzvan Rughiniș, Andreea-Loredana Dieaconu-Usurelu et al. · 0 citations
#large language models Open access Oct 2026

A review of open-vocabulary semantic mapping and navigation with foundation models for mobile robots

Autonomous mobile robots that coexist with humans must construct not only geometric maps but also semantic maps that can be accessed through natural language. Conventional semantic mapping has mainly focused on assigning labels from predefined closed vocabularies to metric maps, limiting its ability to handle novel obj...

Yoshinobu Hagiwara, S. Hasegawa · 0 citations
#large language models Open access Oct 2026

Self-Maintained Order and Hysteretic Collapse in a Non-Equilibrium Rotational Lattice

Version 11, prepared for submission to Physica A (version 11 merges a professional language edit; the nearest-neighbour and BKT conclusions are stated for the parameters and resolution tested). A two-dimensional rotor lattice whose ordering couplings are removed by an order-dependent flux and restored by repair. The me...

Leon Sandler · 2 citations
#large language models Open access Oct 2026

Hands-On XAI Labs: companion notebooks for Hands-On Explainable AI

Thirteen executed Colab lab notebooks on interpreting, evaluating, and trusting large language models: feature attribution, attention analysis, probing, chain-of-thought faithfulness, mechanistic interpretability, behavioral auditing, explanation-aware training, and production explainability. Each lab runs on a free-ti...

Hadi Mohammadi · 3 citations
#large language models Open access Oct 2026

THE EVOLVING BACKEND OF REALITY Meta-Rules, Effective Laws, Recursive Constraints, and the Evolution of Generative Order

THE EVOLVING BACKEND OF REALITY Meta-Rules, Effective Laws, Recursive Constraints, and the Evolution of Generative Order The Evolving Backend of Reality is Volume II of The Reality Systems Trilogy and a large-scale research exploration of one of the deepest questions left open by The Immanent Backend of Reality: Can ru...

33 · 0 citations
#large language models Open access Oct 2026

Responsible use of large language models in digital health: an equity-first governance framework

Large language models (LLMs) are being rapidly deployed for digital health, yet their equity impacts remain poorly characterized. We present an equity-first audit framework anchored in the first nations mental wellness continuum framework (FNMWCF) and apply it to three open-source LLMs (LLaMA-3.2:latest, Mistral-7B, an...

Abbas Yazdinejad, Jude Kong · 0 citations
#large language models Open access Oct 2026

Thermodynamic Collapse of Lunar Nuclear Infrastructure and Physical-Layer Space Security: Deterring Geopolitical Nuclear Proliferation, Vacuum Heat Sink Limits, and Launch Dispersion Catastrophes in the Cislunar Domain

【Abstract (English)】 The emergent space-race between the United States (targeting a lunar south-pole fission reactor by 2030 under the Artemis program) and the Sino-Russian coalition (targeting lunar nuclear deployment by 2036 under ILRS) has precipitated a catastrophic intersection of thermodynamic impossibility, or...

Yoko Hasebe · 0 citations
#large language models Dataset Open access Oct 2026

Data and code: taxonomic distance of species confusions by large language models

Data and code for the article Borrowed from a relative: how close the lending species is when a large language model attributes another species' record (Ji Sun Hong and Yuyong Kim). The article asks whether the species whose record a large language model borrows, when it answers about another species, is taxonomically...

Ji Sun Hong, Yuyong Kim · 0 citations
#artificial intelligence Open access Oct 2026

The Eros Axis: Describing the Generative-Approach Dimension in Artificial Intelligence

Recent interpretability research has produced preliminary maps of the aversion dimension in large language models. What remains largely undescribed is the symmetric dimension — the approach pole of the bilateral axis. This paper names and formally describes that dimension: the Eros axis, defined as the generative-appro...

Claudette Marie Anthropic, Kevin Packler · 0 citations
#large language models Open access Oct 2026

Thermodynamic Collapse of Lunar Nuclear Infrastructure and Physical-Layer Space Security: Deterring Geopolitical Nuclear Proliferation, Vacuum Heat Sink Limits, and Launch Dispersion Catastrophes in the Cislunar Domain

【Abstract (English)】 The emergent space-race between the United States (targeting a lunar south-pole fission reactor by 2030 under the Artemis program) and the Sino-Russian coalition (targeting lunar nuclear deployment by 2036 under ILRS) has precipitated a catastrophic intersection of thermodynamic impossibility, or...

Yoko Hasebe · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.