Skip to content

Author

Tadachika Ozono

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access 2026

Navigating the Certainty Gap: The Hydraulic Effect of Evidence-Constrained LLMs in Professional Decision-Making

Large Language Models (LLMs) are increasingly deployed in high-stakes environments, including infrastructure auditing, medical assessment, legal analysis, and peer review. However, these failures arise less from knowledge limitations than from systematic misinterpretation of modality (fact vs. possibility) and unstable decision commitment, leading to two opposing failure modes: false confirmations (over-commitment) and over-abstention (excessive conservatism). To address this, we propose the Evidence–Constrained (EC) Framework, a structured paradigm for regulating LLM decision boundaries through two complementary mechanisms: a Modality Filter (“the Brake”), which enforces strict separation between factual and conditional evidence, and a Normalcy Rule (“the Accelerator”), which enables inference of non-events from routine reporting structures. These terms are used as analytical metaphors describing opposing influences on commitment behavior rather than literal computational operators. We evaluate the framework through a cross-domain experimental study spanning construction, medical, legal, and peer-review tasks, compared against domain-informed human evaluation. Results reveal a systematic trade-off phenomenon, termed “Hydraulic Effect”, where reducing false confirmations increases abstention, reaching up to 35%. We further identify recurring cross-domain failure patterns, including temporal-modality confusion, lexical overheating, and the Condition Precedent Gap. Rather than yielding a single configuration that is uniformly optimal across the evaluated benchmark, the EC framework exposes a persistent trade-off between decision safety and coverage, suggesting that reliable LLM deployment may require explicit calibration of commitment thresholds under uncertainty. These findings provide exploratory guidance for the design of uncertainty-aware LLM workflows. The reported results should be interpreted as evidence from controlled behavioral experiments on carefully selected ambiguity-sensitive benchmark cases rather than as estimates of real-world deployment performance.

Florence Gundidza, Masato Kikuchi, Tadachika Ozono · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.