Skip to content

Author

Shane Culpepper

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

GRAFT: Graph-Matched Retrieval and Fusion of Tables in Data Lakes

Autonomous data agents resolve analytical queries by retrieving and reasoning over evidence in tabular data lakes. Existing methods score tables independently against the query and ignore the joinability and unionability that link them, returning fragmented evidence that downstream agents cannot integrate. We propose GRAFT (Graph-matched Retrieval and Fusion of Tables), structured around two principal contributions. First, we cast table retrieval as a graph matching problem between a query-derived intent graph and a heterogeneous data lake graph, and introduce IGMS, a log-determinant reward that couples semantic relevance, structural compatibility, and evidence diversity in a single objective. Second, we recast subgraph generation as a Markov decision process and learn a value function via implicit Q-learning on self-generated trajectories produced by a canonical compression operator that inverts the homomorphism. We further design a three-stage online pipeline that exploits anchor reachability, predicate admissibility, and reward monotonicity to greatly prune the candidate space before exact IGMS evaluation. On Spider and BIRD adapted to the tabular data lake setting, GRAFT achieves the best Recall, Precision, F1, and Sufficiency among point-wise, greedy-expansion, and structure-aware baselines, with relative gains of 7.8% in F1 and 10.6% in Sufficiency over the strongest baseline, while maintaining high search efficiency.

Daomin Ji, Hui Luo, Zhifeng Bao et al. · 0 citations
Book Open access Jul 2026

SynthIR: The First Workshop on Synthetic Content in Information Retrieval Ecosystems

The proliferation of AI-generated content is fundamentally altering the information ecosystems in which retrieval systems operate. Search engines, recommender systems, and retrieval-augmented generation pipelines increasingly function in mixed information environments where synthetic and human-authored content are tightly interwoven, raising system-level challenges for information retrieval. Key issues include limitations in evaluation validity, as traditional metrics designed for human-authored corpora fail to capture the distinctive properties of AI-generated content; shifts in retrieval behavior and ranking dynamics, as systems may inadvertently favor procedurally generated but weakly grounded information; and challenges to user trust, as assumptions about the provenance and reliability of retrieved human content become more difficult to distinguish from generated content. Rather than focusing only on model-centric performance comparisons, this workshop aims to provide a forum to analyze these implications with an emphasis on reflection, evaluation, and human-centered system design, and to foster community-driven discussion that may inform future evaluation efforts, including potential shared tasks or tracks in venues such as TREC, CLEF, FIRE, or NTCIR.

Ping Liu, Zhedong Zheng, Shane Culpepper et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.