CORE-LLM-Bench is a controlled neurosymbolic benchmark for evaluating ontology-grounded reasoning in large language models. Version 1.1.1 is a corrective release that supersedes v1.1.0 for future use. It contains 9,048 unique question-hop instances: 6,032 binary question-answering (BQA) instances and 3,016 open-ended q...
Julie Loesch, Sara Falahatkar, Nicole Kilk et al.· Zenodo (CERN European Organi...· 0 citations
A configurable remote judge for Kayenta that produces a canary verdict from a large language model or a vision-language model, three explicit hybrid policies that combine it with the genuine NetflixACAJudge verdict, a tuned statistical ensemble baseline, and a seeded nine-family fault-injection benchmark of 180 labelle...
Vlad-Stefan Dieaconu, Răzvan Rughiniș, Andreea-Loredana Dieaconu-Usurelu et al.· Zenodo (CERN European Organi...· 0 citations
Every result artefact behind the paper, in the directory layout the analysis code expects, together with the scripts that recompute the paper's tables from them and the analysis code that turns them into every number the article reports. The testbed that produced the measurements is the kayenta-ai-canary-judge reposito...
Vlad-Stefan Dieaconu, Răzvan Rughiniș, Andreea-Loredana Dieaconu-Usurelu et al.· Zenodo (CERN European Organi...· 0 citations
Autonomous mobile robots that coexist with humans must construct not only geometric maps but also semantic maps that can be accessed through natural language. Conventional semantic mapping has mainly focused on assigning labels from predefined closed vocabularies to metric maps, limiting its ability to handle novel obj...
Yoshinobu Hagiwara, S. Hasegawa· Advanced Robotics· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Version 11, prepared for submission to Physica A (version 11 merges a professional language edit; the nearest-neighbour and BKT conclusions are stated for the parameters and resolution tested). A two-dimensional rotor lattice whose ordering couplings are removed by an order-dependent flux and restored by repair. The me...
Leon Sandler· Zenodo (CERN European Organi...· 2 citations
Thirteen executed Colab lab notebooks on interpreting, evaluating, and trusting large language models: feature attribution, attention analysis, probing, chain-of-thought faithfulness, mechanistic interpretability, behavioral auditing, explanation-aware training, and production explainability. Each lab runs on a free-ti...
Hadi Mohammadi· Zenodo (CERN European Organi...· 3 citations
THE EVOLVING BACKEND OF REALITY Meta-Rules, Effective Laws, Recursive Constraints, and the Evolution of Generative Order The Evolving Backend of Reality is Volume II of The Reality Systems Trilogy and a large-scale research exploration of one of the deepest questions left open by The Immanent Backend of Reality: Can ru...
Large language models (LLMs) are being rapidly deployed for digital health, yet their equity impacts remain poorly characterized. We present an equity-first audit framework anchored in the first nations mental wellness continuum framework (FNMWCF) and apply it to three open-source LLMs (LLaMA-3.2:latest, Mistral-7B, an...
Abbas Yazdinejad, Jude Kong· Scientific Reports· 0 citations
【Abstract (English)】 The emergent space-race between the United States (targeting a lunar south-pole fission reactor by 2030 under the Artemis program) and the Sino-Russian coalition (targeting lunar nuclear deployment by 2036 under ILRS) has precipitated a catastrophic intersection of thermodynamic impossibility, or...
Yoko Hasebe· Zenodo (CERN European Organi...· 0 citations
Data and code for the article Borrowed from a relative: how close the lending species is when a large language model attributes another species' record (Ji Sun Hong and Yuyong Kim). The article asks whether the species whose record a large language model borrows, when it answers about another species, is taxonomically...
Ji Sun Hong, Yuyong Kim· Zenodo (CERN European Organi...· 0 citations
Recent interpretability research has produced preliminary maps of the aversion dimension in large language models. What remains largely undescribed is the symmetric dimension — the approach pole of the bilateral axis. This paper names and formally describes that dimension: the Eros axis, defined as the generative-appro...
Claudette Marie Anthropic, Kevin Packler· Zenodo (CERN European Organi...· 0 citations
【Abstract (English)】 The emergent space-race between the United States (targeting a lunar south-pole fission reactor by 2030 under the Artemis program) and the Sino-Russian coalition (targeting lunar nuclear deployment by 2036 under ILRS) has precipitated a catastrophic intersection of thermodynamic impossibility, or...
Yoko Hasebe· Zenodo (CERN European Organi...· 0 citations