Reading With, Not Through, AI: A Human-Centered Framework for Multimodal Literature Education
Abstract
Background: Multimodal generative artificial intelligence (GenAI) can process written language, images, typography, layout, and interface-like elements, creating opportunities for literature education but also new epistemic risks. Vision-capable models may combine accurate recognition, plausible inference, and invented visual detail in a single fluent interpretation. The central challenge is not access to multimodal information, but disciplined separation of observation, inference, interpretation, and human judgment. Objective: This perspective paper develops Multimodal Evidence-Bounded Interpretation (MEBI), a human-centered framework designed to preserve close reading, interpretive plurality, evidential accountability, and student authorship when multimodal GenAI is used with literary texts. Methods: The study uses an evidence-informed conceptual synthesis and a worked multimodal case. Recent peer-reviewed research on GenAI-supported learning, student agency, literary analysis, multimodal AI, and visual hallucination is integrated with established scholarship on multimodality and graphic narrative and with human-centered AI guidance. The framework is demonstrated through selected scan locations from Nora Dåsnes’s Czech graphic narrative Na kočičí svědomí (2024). Results: MEBI comprises five design propositions—human-first encounter, modal traceability, observation–inference separation, preserved interpretive plurality, and human adjudication and authorship—and a five-phase sequence: Encounter, Describe, Hypothesize, Audit, and Adjudicate & Author. Five evidence tags track verbal, image, sequential-spatial, typographic-graphic, and digital-interface evidence. A worked case, classroom sequence, evidence ledger, and assessment heuristic operationalize the framework. Conclusion: MEBI positions AI as a generator of descriptions, hypotheses, counter-readings, and questions rather than as interpretive authority. It offers a researchable design for multimodal literature education while requiring empirical validation across genres, languages, age groups, and educational contexts.