Skip to content

Author

Lavdim Menxhiqi

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Hallucination Mitigation in Large Language Model-Based Tool Recommendation: A Cross-Provider Architectural Ablation Study Across Two Model Generations

In a closed-inventory large language model (LLM) system such as Online-CADCOM, which recommends engineering tools from a verified inventory, we measure inventory non-compliance, that is, a mention-level event in which the model recommends a tool not present in the verified inventory. We use this inventory-relative sense of hallucination throughout: an out-of-inventory mention may be a fabricated tool or a real commercial tool absent from the curated inventory, so the metric reports inventory non-compliance rather than factual fabrication. We evaluate a three-mechanism mitigation stack consisting of database-grounded context injection, fixed vocabulary constraints, and enforced JavaScript Object Notation (JSON) output across three commercial LLM providers (OpenAI, Anthropic, Google), two model generations, and two output modes (standard and reasoning), totaling 6912 Application Programming Interface (API) calls over 12 configurations. Under a recall-equalized detector adopted as the primary metric, the inventory non-compliance rate, which we denote the hallucination rate (HR) following common usage, decreases from roughly 69–80% to 4–13% under the full architecture. The cross-provider average is similar across the two generations tested (8.5% Generation 1 (Gen1), 6.9% Generation 2 (Gen2)), although per-provider directions diverge. We also examine the C3 configuration, in which only JSON output enforcement is active without grounding. A naive detector reports a large hallucination increase over the unconstrained baseline (+10.1 percentage points (pp) Gen1, +15.1 pp Gen2), but we show this gap is largely a detection-format artifact: structured JSON fields make out-of-inventory tools easy to extract, whereas the same real tools are frequently missed in free text. Under a recall-equalized detector the gap narrows to +2.6 pp (Gen1) and +4.8 pp (Gen2) and remains statistically significant only for two current-generation models, indicating a small, current-generation effect rather than a universal one. Reasoning-mode models provide no statistically significant improvement under architectural constraints. A frequency-weighted audit shows that the majority of remaining out-of-inventory mentions correspond to real engineering tools absent from the platform’s inventory. Under the full architecture, roughly half of responses (pooled Pany≈49.5%) still contain at least one such mention, indicating that handling unseen tools remains an open challenge for closed-inventory recommendation systems. Our evidence comes from a single engineering platform with four related electronic-design and power-electronics domains, so the findings characterize this setting rather than recommendation domains in general.

Lavdim Menxhiqi, Galia Marinova · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.