Skip to content
Book Open access

Towards Next Graph Token Prediction: Discrete Graph Tokenization for Structural Reasoning in Large Language Models

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · 0 citations · 47 references

Abstract

Large Language Models (LLMs), empowered by autoregressive next-token prediction, have demonstrated strong reasoning capabilities. Extending this paradigm to graph data requires next graph token prediction, yet existing graph tokens struggle to balance two competing requirements: capturing higher-order structural semantics with high information density, and remaining strictly reversible for faithful decoding. Concretely, first-order textual encodings are reversible but long and semantically sparse, while continuous encodings capture high-level semantics but inevitably lose exact topology. To this end, we propose GraphVulcan, a novel framework that enables discrete reversible and semantic-rich tokenization of graphs using a vocabulary of canonical graphlets. Our approach preserves full structural fidelity while enabling LLMs to natively reason over compositional graphlets through graph-level next-token prediction. We propose a three-stage training paradigm: (1) Structural Semantic Pretraining to learn graph token compositionality, (2) Multi-task Fine-tuning on large-scale CoT-augmented reasoning examples, and (3) Reinforcement Learning to explore and refine reasoning paths. Experiments show that GraphVulcan outperforms first-order encoding baselines across 7 graph reasoning tasks and 3 real-world benchmarks while achieving higher computational efficiency. Our work demonstrates that structure-aware discrete tokenization is a feasible way toward general-purpose graph–language models capable of structural reasoning. Codes are available at: https://github.com/alibaba-behavioral-risk-control/GraphVulcan

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.