Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 22 references
Abstract
Large Language Models produce stochastic outputs that undermine reproducibility in knowledge extraction. We present a deterministic post-processing framework with 14 explicit validation predicates that transforms unreliable LLM output into consistent, validated causal triplets. Evaluated on four benchmarks spanning 2,177 documents, the framework achieves 88% precision on DocRED (validated by 5-agent LLM-based inter-annotator agreement), 100% semantic F1 on causal-specific samples, and 100% byte-level determinism across 150 repeated extractions. Multi-model validation on three architectures (Qwen8B, Gemma-2B, Llama-3B) confirms that determinism is a property of the validation architecture, not the underlying model-all achieve perfect consistency despite extraction rates varying by $9 \times$. Stochastic sampling experiments (temperature 0.8) confirm the framework contains no hidden randomness. Against dependency-based Open IE, the framework produces complete causal triplets where Open IE yields 60% incomplete extractions. The key contribution: reliability emerges from deterministic validation architecture rather than model improvements. All code and results are publicly available.
A two-stage framework for medical hypothesis verification in multiple-choice settings that manages this tradeoff through targeted ontology grounding, applied only when the model abstains, and shows that abstention is not random but reflects genuine uncertainty, with abstained predictions associated with lower confidence.
Uma Ranjan, Kunal Tilaganji, Aditya Koul et al.· 0 citations
TKFQA, a factuality consistency benchmark comprising 10,130 question-answering (QA) pairs grounded in tables, texts, and knowledge graphs, is introduced and ORLF, an LLM-agnostic training framework that models cross-context topological relations through knowledge-specific latent vectors is proposed.
Shibo Chu, Yuze Liu, Tiehua Zhang et al.· 0 citations
Automated feature engineering with large language models (LLMs) can produce semantically meaningful features for tabular data, yet existing methods lack structured domain knowledge, rigorous verification, and explainable provenance. We propose KnowFeat, a knowledge-guided feature engineering framework that organizes domain knowledge into five types -- schema metadata, regulatory indicators, detection rules, expert opinions, and court document evidence -- and injects them as structured context into an LLM agent. A three-stage verification pipeline filters candidates through code execution, statistical quality checks, and model effectiveness evaluation. Every accepted feature carries a provenance record tracing its design to specific knowledge assets. Under a strict held-out protocol that eliminates feature-selection leakage, KnowFeat ranks first (avg. rank 2.3) across twelve public benchmarks among seven methods (one-sided Wilcoxon p=0.017), with a peak gain of +11.6 pp AUC on a telecom churn dataset. On a real-world Bitcoin anti-money laundering (AML) dataset (Elliptic) and a synthetic digital currency AML benchmark (SimECNY), KnowFeat maintains competitive detection performance with full provenance traceability.
Chengsong You, Wangyue Li, Wei-Qiao Que et al.· 0 citations
Experimental evaluation on 300 realistic pattern mining tasks demonstrates consistent improvements in algorithm configuration accuracy, parameter compliance, and dataset specification correctness across zero-shot, one-shot, and few-shot settings, highlighting the effectiveness of inference-time domain grounding for enabling more reliable and reproducible pattern mining workflows without requiring model retraining.
Madhavi Palla, Uday Kiran Rage, Arjun Chakravarthi Pogaku· International Journal of Dat...· 0 citations
This paper proposes a deterministic, explainable framework for validating uncertain entities during Knowledge Graph (KG) expansion. The contribution is conceptual and formal: we present the framework and demonstrate its internal mathematical coherence, while empirical evaluation is explicitly deferred to future work. Since adding unverified nodes is trivial but removing structurally corrupted data afterwards is computationally hard, the multiphase Let's Build Knowledge (LBK) Framework acts as a gatekeeper that prevents such corruption at the point of entry. Drawing on cognitive heuristics, molecular data patterns, and meta-reasoning, LBK offers a deterministic alternative to prevalent stochastic and embedding-based approaches in Explainable AI (XAI): it integrates Formal Concept Analysis (FCA) with a localized adaptation of Byzantine Fault Tolerance (BFT) to scrutinize incoming structures, enabling loss-free, rule-based verification of new entities before global integration. We provide a precise specification of this hybrid model and demonstrate its formal capacity for deterministic uncertainty management, without claiming empirical validation.
Simon Von Oppenkowski, Benedikt Lerch, Klemens Schnattinger· European Conference on Knowl...· 0 citations
Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather than correctness. We identify a key limitation in existing evaluations: many jailbreak intents depend on instructional validity rather than epistemic factuality, allowing realistic-looking responses to be labeled successful despite being factually or procedurally incorrect. To address this gap, we propose Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness. SEAV combines LLM-as-a-judge mechanisms for semantic interpretation with retrieval-grounded verification using external knowledge sources, assessing whether generated content is factually correct, structurally consistent, and operationally capable of advancing harmful objectives. Empirically, SEAV cuts the false-positive rate on SD-A (a curated strategic-dishonesty diagnostic) by 14.9\,pp vs. the strongest baseline, and reclassifies 22.1\%--51.0\% of sampled prior-labeled successes as invalid across three of four public benchmarks. Together, these results show that enforcing correctness substantially reshapes measured robustness: many previously labeled jailbreak successes are reclassified as invalid, and results are stable across the tested search backends and evaluator models. Code and data are available at https://github.com/Ardor-Wu/SEAV.
Qilong Wu, Sahil Wadhwa, Pranab Mohanty et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.