Semantic Parsing for Evaluating Large Language Models: Separating Linguistic Abilities with YARN
A layer-wise analysis indicates that surface-level features such as temporality and negation are captured more reliably than deeper semantic phenomena like quantification in large language models, highlighting the limited capacity of current LLMs to generate fully formal meaning representations.