Sep 2026· International Journal For Multidisciplinary Research· 0 citations· 21 references
TL;DR
An Adaptive Intelligence Router is set out that chooses among deterministic programs, retrieval and service interfaces, small language models, more capable language models, more capable language models, and accountable human review.
Abstract
Which solver should handle a given task — with which tools, how much verification, and what opportunity to escalate? Every AI application answers that question, and most answer it implicitly. Defaulting to a large language model spends reasoning on work that never needed it. Choosing a small model on price alone does not remove the cost; it relocates the cost into errors and rework. This paper sets out an Adaptive Intelligence Router that chooses among deterministic programs, retrieval and service interfaces, small language models, more capable language models, and accountable human review. A task’s demands are carried by a complexity profile in which every component has a defined role: three dimensions generate the admissible set, one fixes the reliability floor and the mandatory oversight level, and the rest are predictive features and nothing more. Selection then becomes a constrained sequential decision problem spanning routing, verification, retries and handoffs. I distinguish the pre-execution success probability that chooses an action from the post-verification probability that licenses accepting its output, and show that only the second supports a bound on residual error among accepted work. An analytical illustration then parameterises a two-stage cascade by verifier recall and false-flag rate instead of escalation fraction. Doing so exposes a cost-quality frontier whose cheapest point is also its least accurate. The evaluation protocol proposed here compares static routing, fixed cascades, learned routing, decomposition, and a non-routing caching baseline under quality, latency and residual-risk constraints, and it counts the cost of certifying the router itself. The paper reports a design and a protocol; no router was trained and no benchmark results are given.
CAS (causal active sequential experimentation), which targets evaluation to model-workload pairs and repeats the test as evidence accumulates, to ask whether one assignment stays optimal across every quality table consistent with the evidence.
This work presents ToolGate, which turns repeated answer checking and difficulty screening into an auditable process while leaving domain design and final review to experts.
Ke Zhang, Yan-Kang Liu, Roya Zandi et al.· 0 citations
This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...
This paper investigates whether there exists a token-budget threshold, below which the overhead of planning and verification hurts performance and above which it helps, and evaluates two systems on FinQA and TAT-QA financial reasoning tasks.
Thomas Nolasque, J. Grey, Calista Pham et al.· 0 citations
Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each can...
D. Rajput, Nirdesh Chauhan, S.Rao Kosaraju· 0 citations
Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in...
Tobias Deußer, L. Sparrenberg, R. Sifa· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.