ASH: A Systematic Automated Evaluation Framework for Creative Text Generation in Cuisine Transfer Tasks
Abstract
Large language models (LLMs) have shown promise in creative domains such as the culinary arts, yet evaluating their ability to balance factual fidelity with authentic innovation remains a challenge. We present an extended empirical study of the $\mathcal {ASH}$ (authenticity, sensitivity, harmony) benchmark, designed to evaluate cultural creativity in the cuisine transfer task. While prior work introduced the core metrics, we expand the evaluation framework by integrating systematic prompt engineering strategies and empirically validating them against human judgments. Our study combines both automated and human evaluations across 4,800 recipes generated from 800 standardized prompts spanning 20 base dishes and 40 culturally diverse cuisines. Through comparative experiments involving multiple model families and eight distinct evaluator prompt strategies, we analyze how LLMs generate and judge culturally adapted recipes. Our findings reveal that while models consistently achieve high sensitivity scores, they often fail to preserve authenticity or maintain overall harmony, indicating a tendency toward surface-level cultural adaptation rather than genuine culinary coherence. We further show that specific prompting techniques (e.g., explicitly defining scoring scales) help align LLM evaluators with human annotators, although the top-ranked strategies have overlapping confidence intervals; we therefore present Prompt 3 as a practical operating choice rather than a uniquely optimal one. This work provides a unified framework for studying creativity and cultural adaptation in generative AI. The code and dataset are publicly available at https://github.com/ dmis-lab/ASH2608/