Skip to content

FlavourBench : Ranking Frontier Language Models with Executable Culinary Ground Truth

· 0 citations · 19 references

TL;DR

FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth, is introduced, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth.

View source

Similar papers

Preprint Aug 2026

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.

Ziyue Wang, Aomufei Yuan, Yi-Ran Yao et al. · 0 citations
#machine learning Preprint Aug 2026

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF); the resulting frozen, source- and easiness-balanced weights score models in the other half, while equally short matched-random subtests control for generic subtest variation. Full-benchmark and low-DIF rankings remain strongly correlated ($\tau_b=.900$--$.948$). Yet in four of five benchmarks, 30.9--47.1\% of cross-family pairs initially within one percentage point reverse order, exceeding their matched-random medians by 16.9--28.6 percentage points (all $p=.001$). The fifth benchmark shows no reliable excess ($-0.9$ points, $p=.689$). The pattern survives all pre-specified population perturbations, and residual item--family signatures replicate across owner halves; however, no family shows a consistent advantage across benchmarks. Thus, globally stable rankings can still leave individual near-tie orderings sensitive to benchmark composition, and sub-one-point leaderboard gaps should be accompanied by evidence that the implied ordering is composition-robust.

Qi-Yuan Zheng, Yi-Qu Yang · 0 citations
Review Aug 2026

OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models

OenoBench is introduced, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers, and an LLM-driven pipeline in which language models reformat verified facts and audit the result is introduced.

Nikita Khudov · 0 citations
#artificial intelligence Preprint Sep 2026

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or>=2.0,<3, yet their grasp of version-constraint semantics has never been measured directly. We introduce SemVerBench, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo): 240 machine-checkable items with unique answers, built author-neutrally from four balanced sources (each ecosystem's official test suite plus three frontier LLM proposers) and labeled by a non-circular two-implementation oracle. Evaluating six frontier models, we find systematic, predictable per-mechanism blind spots: a partial-comparator carry rule (>1.2 means>=1.3.0) traps every model on Cargo (near 60%), and although standard PEP 440 prefix matching is universal, on zero-pad/post-release corner cases GPT-5.1 collapses (0/26) while Claude stays at 97-100% (verified on a 67-item oracle-validated set). Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar). The failures look more like an activation/application gap than a knowledge gap: injecting the rule or a light correct hint recovers most errors, whereas interval decomposition does not, and models are at ceiling on the basic forms of the same rules. An author-stratified analysis finds no statistically significant self-favoritism. Because the task is verifiable and a free, 100%-correct resolver exists, tool delegation reaches ~100%: coding agents should delegate version resolution to a resolver rather than reason about versions in-head.

Qi-Bai Chen, Ze-Ming Liu · 0 citations
Jul 2026

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells.

Yan-Shi Li, Xue Bai, Shu-Man Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.