Evaluating a diverse panel of contemporary models under three administration protocols, it is shown that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-biased, and overconfident.
AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation denotes a known formal object. We study this challenge through theorem recognition: given an equivalence-preserving transformation of a theorem condition, a model must recover the theorem identity associated with the standard statement. We introduce TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving formula-level transformations. Rather than paraphrasing theorem text, TREAT changes the mathematical form of theorem conditions themselves, expressing known results through residual equations, witness statements, optimization identities, set relations, operator forms, and proof-intermediate characterizations. Starting from scraped theorem pages, we filter for entries with usable mathematical expression forms, extract canonical theorem conditions, and generate transformed variants with recorded assumptions and inverse mappings. The final corpus contains 737 theorem identities and 29,480 transformed rows. On a test panel, the best model retrieves the correct theorem identity in only 60.73% of cases. Other systems reveal different failure modes, including abstention, wrong detection, and malformed outputs. These suggest that theorem knowledge can be fragile under equivalent changes in representation. TREAT therefore provides a controlled testbed for evaluating representation-robust access to formal knowledge, with broader relevance to domains that require stable target objects, explicit equivalence relations, validation procedures, and auditable scoring.
Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can obscure whether a model can verify a non-canonical but valid solution trace. We introduce a controlled linear-equation benchmark for evaluating LLMs in the evaluator role. Each instance asks the model to judge final-answer correctness, step-level trace correctness, and the first incorrect step. Our evaluation of state-of-the-art open LLMs reveals a significant robustness gap: models that accurately evaluate canonical solutions often fail when presented with perturbed but logically equivalent variants. Across GPT-OSS 20B, Qwen3-14B, and Phi-4-Reasoning, base models perform well on canonical traces but degrade substantially on perturbed traces, especially for error localization. On valid perturbed traces, base-model false-rejection rates reach 75.6-85.3%, showing strong sensitivity to canonical solution form. Supervised fine-tuning, distillation, and test-time compute improve robustness in some settings, but gains are model dependent and can trade off against canonical performance. The results show that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.
Crop modelling is essential for agricultural water management but often relies on simplified water balance routines that limit representation of soil moisture dynamics. To address this limitation, we developed a coupled model that integrates the 1‐D Richards equation, solved using a finite difference method into the FAO AquaCrop. The coupled model was calibrated and validated using soil moisture, canopy cover, above‐ground biomass and seed cotton yield data from field experiments in the southeastern United States. Compared with hourly field measurements of soil moisture, AquaCrop–Richards achieved an average root mean square error (RMSE) of 0.023 m
3
m
−3
across three soil depths over the growing season. Model performance for canopy cover, biomass and yield resulted in RMSE values of 12.18%, 1.77 t ha
−1
and 0.96 t ha
−1
, respectively, against observations. Under fully irrigated conditions, both models produced statistically indistinguishable yield estimates. However, under rainfed conditions, AquaCrop simulated 15.5% higher yields than AquaCrop–Richards. Analysis showed that AquaCrop produced rapid stepwise drainage, resulting in root‐zone water content 33%–37% lower than the coupled model. This reduced soil moisture triggered earlier water stress which led to yield overestimation. These results indicate that AquaCrop‐Richards improves soil moisture representation and is robust under water‐limited conditions.
Krishna Panthi, Vidya Samadi, Carlos Toxtli· Irrigation and Drainage· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.