Skip to content
Preprint

Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

Results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements in reference-free agentic data cleaning.

Abstract

Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.

View source

Similar papers

Jul 2026

Auditing Provenance Sensitivity in LLM Agent Action Selection

A target-specific authorization audit is introduced that labels context factors separately for each tool and argument target and holds the task, proposition, position, and policy fixed while changing only the proposition's source authority.

Jun-Hui Liao · 7 citations
Preprint Aug 2026

How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

On ClimateCheck claim-only and fine-tuned models outperform zero-shot LLM and top-performing AVeriTeC 2025 systems, highlighting that noisy evidence can degrade veracity prediction and replacing retrieved evidence with gold annotations improves veracity accuracy, confirming retrieval remains primary bottleneck.

Aida Usmanova, Z. Iklassov, Markus Leippold et al. · 0 citations
Jul 2026

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

Agent benchmarks should report success together with whether the evaluated information state supported it, and whether the evaluated information state supported it should report success together with whether the evaluated information state supported it.

Jingkun Luo, Dai-Yun Peng · 0 citations
Jul 2026

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

DBA-Bench is presented, a benchmark addressing four gaps between evaluation and production operations: live-environment fidelity, outcome-first evaluation, and controlled scenario reproducibility, which uses instrumented PostgreSQL environments with active workloads, persistent state, and multi-source observations.

Junming Chen, Jun-Yang Jiang, Xu Chen et al. · 0 citations
Preprint Aug 2026

TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis

TrustDABench is introduced, a benchmark that operationalizes two diagnostic questions of LLM reliability and robustness and suggests that stronger evidence-boundary recognition and representation-invariant reasoning are still needed for reliable structured-data analysis.

Boshen Shi, Yize Liu, Chen Zhao et al. · 0 citations
Preprint Aug 2026

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

ABE-Ralph is introduced, a reference-anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8-step workflow, and performs quantitative, qualitative, and code-level verification.

Le-Zhi Yu, Xiaogang Xu, Yuhong Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.