MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research
This paper systematically compares four state-of-the-art LLMs, including Gemini~3.1~Pro, GPT-5.4, and Claude~Opus~4, on a FormalMATH subset and on two real PDE theorems requiring deep domain expertise, evaluating their ability to produce verified Lean~4 proofs and to identify errors in deliberately incorrect proofs.