Benchmark-Driven Selection of AI Can Improve Capabilities of Reasoning Language Models in Epistemology
Abstract Evaluation of reasoning language models gained importance after it was observed that they can combine their existing capabilities into novel traces of intermediate steps before task completion and that the traces can sometimes help them to generalize better than past models. We show that better performance res...