Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer largely from memory, which is least reliable for less promi...
Mohamed Chenene, C. Rosas Hinostroza, A. Stasenko et al.· 0 citations
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to...
Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois et al.· 0 citations
BenCheck is proposed, a package for benchmark validity analysis that encapsulates the main checks performed in a case study on HellaSwag and can be used to audit commonsense reasoning benchmarks, including PIQA, Global PIQA, and Winogrande.
Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar et al.· 9 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.