Seeing the knowledge graph: a multimodal vision–language framework with active learning for cold–start course recommendation
ALARM is proposed, a multimodal active–learning framework pairing a frozen large language model (LLM) that encodes course and learner text, a vision transformer (ViT) that “observes” the KG by reading its rendered subgraph image, and a domain–specific acquisition function that selects which labels to query.