TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
This work introduces TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation, and benchmarks the verifier against human-expert proof judgements on a set of target statements and generated proofs pairs.
Vincent Cohen-Addad, Dimitris Paparas, Ernest van Wijland et al.
· 1 citation