Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what human fact-checkers...
Rui-Hong Zeng, Jonathan Tonglet, Preslav Nakov et al.· 0 citations
Co-FactChecker is proposed, a framework for human-AI collaborative claim verification that translates expert feedback into trace-edits that introduce targeted modifications to the trace, sidestepping the shortcomings of dialogue-based interaction.
It is argued that human participation may persist even with highly capable AI systems for three distinct reasons, and this perspective has important implications for the limits of automation and for the design, evaluation, and ethics of future AI systems.
Reviewer comments naturally relate to specific parts of the reviewed paper, yet grounding these comments to the underlying evidence is difficult due to long multimodal documents. Existing benchmarks do not capture this setting and largely focus on explicit, information-seeking queries. We introduce ReGround, a large-sc...
This work proposes a new intrinsic reward for learning tool use in MLLMs without the need for additional labels or warm-start SFT, and finds that recall is the metric that most strongly correlates the zoom-in region with final task performance.
ChartAttack is presented, a framework for evaluating how MLLMs use design misleaders to generate charts that induce incorrect interpretations and AttackViz is introduced, a chart question-answering (QA) dataset labeled with effective misleaders and their induced incorrect answers.
Jesús-Germán Ortiz-Barajas, Jonathan Tonglet, Vivek Gupta et al.· arXiv.org· 0 citations
It is established that a prompting strategy that scores multiple aspects of the writing together is the most effective, paving the way for more effective classroom deployment of modern LLMs.
Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov et al.· 2 citations
This work evaluates the performance of a news article retrieval pipeline, NewsRECON, which leverages a corpus of over 85,000 articles and investigates the potential of news article corpora as an alternative to RIS, linking images to relevant articles to infer their dates and locations from article metadata.
Jonathan Tonglet, Iryna Gurevych, T. Tuytelaars et al.· arXiv.org· 1 citation
This work proposes CORE-T, a scalable, training-free framework that enriches tables with LLM-generated purpose metadata and pre-computes a lightweight table-compatibility cache, and uses 1.20x fewer total selection tokens than LLM-intensive baselines.
Hassan Soliman, Vivek Gupta, Dan Roth et al.· arXiv.org· 2 citations· ⚡1
Murano is an open source framework for designing, running, and reproducing mechanistic interpretability studies of large language models, intended for researchers across disciplines and builds on existing interpretability and machine learning libraries.
Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang et al.· 0 citations
OctoLong is introduced, a context engineering pipeline that instruments an AST parser, a language server backend, and a package manager to facilitate the recursive retrieval of code references, enabling the curation of dependency-rich code contexts of millions of tokens in length.
Indraneil Paul, F. Helm, Goran Glavas et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.