Tabular Anomaly Detection (TAD) plays a fundamental role in securing real-world applications. Despite rapid advances in TAD, the prohibitive cost of human-centric label annotation remains a primary bottleneck for large-scale production systems. To alleviate this bottleneck, we propose a novel ''coarse-to-fine'' label annotation pipeline to improve labor efficiency through a coarse-grained label annotation and fine-grained human verification. Specifically, Large Language Models (LLMs), with their strong cross-domain capabilities, serve as a promising solution for the coarse-grained annotation stage. However, effectively generalizing LLMs to coarse-grained annotation remains challenging due to the inability to ground semantic priors in rigorous deduction, as well as the overfitting risks inherent in single-domain fine-tuning. Accordingly, we introduce TaDGeneral, a large-scale cross-domain corpus constructed by fusing deductive reasoning paths from diverse domains. This design bridges the reasoning gap while preventing the memorization of local shortcuts. Building upon this, we develop TaDFM, a foundation model tailored to internalize generalizable deductive logic for effective zero-shot annotation. Extensive experiments on both public and large-scale real-world TAD datasets demonstrate the superiority of TaDFM over representative methods, with its practical value further validated by an industrial case study. Code: https://github.com/cshhzhao/TaDFM.
Haihong Zhao, Aochuan Chen, Miao Peng et al.· Proceedings of the 32nd ACM...· 0 citations
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce $R^3$-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.
Peisong Wang, Zhiwei Ma, Bo-Wen Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.