Modern cloud platforms face escalating diagnostic challenges due to scale-driven emergent behaviors and intricate system interdependencies. However, contemporary diagnostic systems can only handle well-understood failure patterns, leaving two critical blind spots that impede system reliability: (1) <italic>Long-tail di...
Zhu Chang, Yin-Jun Wu, H. Feng et al.· IEEE Transactions on Knowled...· 0 citations
This work reframe schema linking as uncertainty-aware schema-need inference over multiple plausible SQL paths, where the system distinguishes required schema items from path-dependent uncertain ones and acquires evidence only where needed.
Huawei Zheng, Sen Yang, Zhaorui Yang et al.· arXiv.org· 2 citations
This work proposes a policy-centric training paradigm that reframes skills as a dynamic training scaffold and converts rollout groups from the latest policy into evidence cards and uses task-specific evaluation to adjust the context used in subsequent rollouts.
Yipeng Shi, Zhi-Peng Ma, Yue Wang et al.· arXiv.org· 0 citations
DataClawEval is introduced, the first comprehensive benchmark designed specifically to evaluate the end-to-end task completion capabilities of autonomous agents in real-world data engineering scenarios, and it comprises 100 rigorous, end-to-end tasks spanning five execution engines.