VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation, and shows that strong label-level performance can conceal substantial errors in complete moderation records.
Mingyu Yuan, Shengtao Wen, Lingbing Guo et al.· 0 citations
Praxist is introduced, a lineage-centered generational system that converts reproducible artifacts and evaluator outcomes into a typed evidence graph of findings, lane-structured frontiers, and agendas, and Separating local artifact construction from cohort-level evidence synthesis lets later attempts inherit validated mechanisms, unresolved claims, and useful constraints.
Jin Li, Ahmed Murtadha, Zhiying Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.