Skip to content
#small language model Open access

Concept Tree Learner (CTL): An Incremental and Interpretable Symbolic Framework for Binary String Rule Induction

Aug 2026 · Applied Sciences · 0 citations · 23 references

TL;DR

The Concept Tree Learner is introduced, an incremental and interpretable symbolic framework that induces logical concepts over binary strings from minimal labeled data and generalizes substantially better to unseen strings than both a classical entropy-based decision tree and the RIPPER rule learner.

Abstract

Symbolic and rule-based learning offers a transparent, sample-efficient alternative to the statistical paradigm that dominates contemporary machine learning. Whereas large language models and deep networks approximate target functions from massive corpora without exposing the rules they rely on, many practical problems instead call for compact, human-readable concept definitions learned from only a handful of examples. In this paper, we introduce the Concept Tree Learner (CTL), an incremental and interpretable symbolic framework that induces logical concepts over binary strings from minimal labeled data. CTL represents knowledge as a rooted tree of logical predicates; it extends this tree incrementally as each labeled example is processed and, after observing the data, distills the simplest rule set consistent with all examples through a set-cover filtering step, in accordance with Occam’s Razor. We give formal definitions for the tree structure, its construction and pruning operations, and the rule-selection objective, and we analyze the worst-case time and space complexity of the procedure. A prototype implementation, operating over a small set of atomic binary-string predicates whose numeric arguments scale dynamically with the input length, is evaluated across 29 concept-learning tasks of varying complexity. CTL recovers a consistent concept for every task—the intended one on 26 of the 29 datasets, and an equally consistent alternative on the three whose training set does not uniquely determine it—typically converging before all training examples are exhausted, and produces fully interpretable rule sets. On the same atomic vocabulary, it generalizes substantially better to unseen strings than both a classical entropy-based decision tree and the RIPPER (Repeated Incremental Pruning to Produce Error Reduction) rule learner (95.5% mean accuracy across all tasks—and 100% on the 26 tasks whose training set uniquely determines the target—versus 72.1% and 77.2% respectively), while using fewer and shorter rules. We position CTL within the literature on symbolic machine learning, inductive logic programming, and decision-tree induction, discuss the current limitations of the prototype—its restriction to noise-free binary input, its batch (sorted) training regime, and its fixed predicate vocabulary—and outline concrete directions for extending it toward a self-expanding, hierarchical concept-learning system.

Read PDF

Similar papers

Review Open access Sep 2026

Enduring Strengths of Knowledge-Based Methods in the Era of Machine Learning and Large Language Models: A Task-and-Architecture Perspective

Large language models have shifted AI toward statistical learning, but knowledge-based methods remain essential for tasks governed by combinatorial structure, declarative correctness, strong domain priors, and auditable reasoning. This paper treats the issue as one of task–architecture fit rather than paradigm competit...

Maikel Leon · 0 citations
#natural language process... Preprint Sep 2026

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

A declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning, and links evaluation and learning through a shared semantics.

Kyle Richardson, C. Anderson, Pranav Balakrishnan et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, ra...

Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi · 0 citations

INDUCING KNOWLEDGE

Unknown authors · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.