Back to feed
Review Open access

Hallucination Is Not One Thing: A Two-Axis Taxonomy for Structured Diagnosis in Generative AI

2026 · IEEE Access · Vol 14, pp. 119273-119304 · 0 citations · 45 references

Abstract

Hallucinations, fluent outputs from generative Artificial Intelligence (AI) that are factually incorrect or unfaithful to their inputs, erode model reliability and alignment in high-stakes domains. Although widely recognized, the existing literature still lacks a compact and task-agnostic framework for detecting and mitigating these errors. This paper presents a concise two-axis framework that integrates an “intrinsic-extrinsic” distinction in source attribution introduced by Ji et al. with a “faithfulness-factuality” distinction in contextual grounding surveyed by Huang et al. and Maynez et al., integrating these established axes into a unified four-quadrant structure, yielding four clearly defined hallucination types applicable across tasks, modalities and architectures. This paper shows how this framework reorganizes existing benchmarks, guides detector and mitigator selection, and supports a fine-grained annotation schema. As the primary empirical validation, a case study on the HalluMix benchmark (6,500 naturalistic examples) evaluates two detectors, Claude Sonnet 4.6 and GPT 5.4, across four prompting conditions, and in every case the taxonomy-guided prompt attains the highest detection F1-score. It raises F1 over a minimal generic prompt from 0.8620 to 0.8741 (95% CI [0.8660, 0.8822]) on Claude Sonnet 4.6 ( $p \lt 0.001$ ) and from 0.8800 to 0.9006 (95% CI [0.8928, 0.9081]) on GPT 5.4 ( $p \lt 0.0001$ ). The margin over the strongest non-taxonomic baseline is +0.10 pp on Claude Sonnet 4.6 (not statistically significant; $p = 0.38$ ) and +1.29 pp on GPT 5.4 ( $p \lt 0.0001$ ), indicating that much of the binary detection benefit comes from structured prompting in general, while taxonomy-specific definitions add a further significant increment when model headroom remains and uniquely provide subtype-level diagnostic labels in both cases. To enable subtype-level diagnostic analysis not possible on HalluMix, we conduct a complementary study on a controlled synthetic dataset (2,000 examples). This study exhibits the same ordering on both models—taxonomy-guided F1 of 0.9727 (Claude Sonnet 4.6) and 0.9753 (GPT 5.4), the highest in each case—while additionally producing subtype-level diagnostic labels that the binary generic prompts cannot, and revealing systematic misclassification patterns across quadrants that differ markedly between the two models. Across both naturalistic and synthetic datasets and both architectures, taxonomy-aligned prompting matches or exceeds the strongest generic baseline while uniquely yielding subtype-level diagnostic labels, though the magnitude of binary detection gains is model- and task-dependent. While the taxonomy is conceptually modality-agnostic, we scope empirical validation to text-based tasks and provide a conceptual mapping to multimodal settings as a foundation for future work. By separating where hallucinations originate from how they violate truth conditions, the taxonomy provides a principled basis for structured evaluation and diagnostic benchmarking in text-based settings, with potential to support governance practices as validation extends to additional domains and modalities, advancing the goal of generative AI systems to remain reliable and aligned with human intent and epistemic standards.

Read PDF