From CSV to Fraud Graphs: An Automated LLM-Guided Pipeline for Graph-Based Fraud Detection and Querying
Abstract
Graph-based fraud detection can capture relational dependencies among suspicious transactions, but constructing graph representations from raw CSV files and querying the resulting graph database still require substantial schema engineering. We propose an end-to-end framework that uses a small language model (SLM) to interpret tabular columns, construct star-topology transaction graphs, train graph neural networks for fraud detection, and support natural-language graph querying through Neo4j. The pipeline classifies columns into identifier, relation, feature, and exclusion roles and applies a leakage-aware preprocessing procedure before graph construction. Fraud detection is performed with established GNN baselines and F-GNN, which is used as an existing frequency-aware backbone rather than a new model contribution. The same schema information is reused by a schema-aware Text2Cypher workflow that combines draft-based schema linking with iterative execution-based self-correction. Experiments on two real-world fraud datasets show that F-GNN achieves the highest AUC and F1-Macro among the evaluated baselines on the automatically constructed graphs. On the Neo4j Text2Cypher benchmark, the complete workflow improves the correct-query rate of Qwen2-7B-Instruct from 13.35% to 18.44% on the full view and from 22.19% to 31.82% on alias-bearing queries, while reducing invalid queries from 13.39% to 4.86% and from 26.19% to 9.51%, respectively.