It is argued for a bridging model with more than one axis of disagreement, and for recruiting raters in the languages the current design reaches least, for recruiting raters in the languages the current design reaches least.
Abstract
Community Notes is X's crowdsourced fact-checking system. A note is published beneath the post it corrects only when raters who usually disagree both rate it helpful, a design called bridging. To apply that rule, the system learns who disagrees with whom from the ratings alone, placing every rater and note on one line, the polarity axis. Every scorer in the production pipeline uses a single axis. Refitting the base model these scorers share on the full public data (212.9M ratings, 2.33M notes, 1.07M raters), we find that one axis is too few. The space is at least two-dimensional. The first axis is left/right politics, while the second, which we interpret as trust in institutions, is largely independent of the first. A held-out test confirms that the second axis improves prediction of unseen ratings, while a third adds little. A second rater dimension learned from one set of topics predicts how raters judge COVID and Ukraine notes excluded from the fit, so it does not merely restate subject matter. Among heavily rated notes that barely divide raters politically, the published share falls from 71.5% to 11.7% as second-axis disagreement grows. A one-axis fit records these notes only as weakly polarised and less helpful; the information that raters at one end of the second axis support them is lost. Authors write notes matching their own position on both axes (r = 0.538 and 0.358), and a small minority of raters cast most ratings (Gini = 0.718). Fewer notes are published in the smallest language communities, but the shortfall is in ratings received, not in how the rule treats them. Keeping ratings per note constant, only Hindi stays below the global rate of 10.85%, and Greek moves from 7.76% to 11.68%. We argue for a bridging model with more than one axis of disagreement, and for recruiting raters in the languages the current design reaches least.
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude, and instances of one model monitoring each other could collude.
It is shown that a widely used dataset's stated assumption of no intra-group annotator variability is false, and it is shown that an LLM judge tracks the crowd pool over the expert pool on all three models the authors test, including one from a different vendor.
The official evaluation of the RecSys Challenge 2026 Music-CRS task can disagree with itself and a single composite of nDCG@20, catalog and lexical diversity, and an LLM judge is found, finding internal conflicts.
Tomoya Terai· Proceedings of the Workshop...· 1 citation· ⚡1
Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we decompose all three tasks of SemEval-2026 Task III Track A in...
Yi-Qun Zhang, Pei-Dong Wang, Zi-Han Wang et al.· 0 citations
Attention dominates token mixing, but it collapses relation formation and flow allocation into a single score-to-flow step. We introduce Relation, which separates them by first organizing pairwise evidence into explicit Self and Exchange relations and deriving information flow afterward. Relation first decides whether...
Conference policies distinguish using Large Language Models (LLMs) to polish one's own review from delegating the critique, but current Artificial Text Detection (ATD) methods largely measure surface form rather than the origin of its content. We instead measure the external information carried by a review: information...
Matthieu Dubois, Pablo Piantanida, François Yvon· 0 citations