Skip to content

Author

J. González-Fraga

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

Quantifying Conceptual Brittleness: An Evaluation of Large Language Models on Biographical and Occupational Inference

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet their ability to maintain stable inference under linguistically varied inputs remains poorly understood. We introduce the concept of conceptual brittleness, defined as the failure of an LLM to sustain inference performance when inputs undergo meaning-preserving transformations such as paraphrasing, anonymization, or metaphorical rephrasing. To quantify this phenomenon, we propose a four-step evaluation methodology of increasing complexity: 1) identifying the boundaries of each model’s entity recognition through a notability recognition task; 2) assessing factual reasoning and conceptual inference with and without biographical context; 3) evaluating biographical inference by requiring models to identify individuals from anonymized and subsequently paraphrased biographies, generated using Llama 3-8B; and 4) testing occupational inference by presenting original and paraphrased definitions of professions, with cross-validation using DeepSeek V3 as an alternative paraphrase generator. All evaluations leverage a cross-verified repository of notable individuals spanning 3500 BC to 2018 AD. We evaluated GPT 4, GPT 4o, Llama 3-70B, Llama 4-Maverick, Gemma 3, DeepSeek V3, and Mistral $8\times 7$ B. Results reveal persistent challenges in gender inference, with GPT 4 and Mistral $8\times 7$ B exhibiting the greatest variability and degraded performance when additional context was provided. GPT 4o achieved the highest accuracy in biographical inference. A significant finding was the contrast between entity-level and concept-level brittleness: biographical inference showed a modest average accuracy decrease under paraphrasing (approximately 7.0% on recognized and 7.5% on unrecognized individuals), whereas occupational inference suffered substantially larger drops of 30.14% (GPT 4o), 31.19% (Llama 4-Maverick), and 37.10% (DeepSeek V3) when comparing original definitions versus paraphrased ones (under paraphrases generated by Llama 3-8B). These findings demonstrate that, although LLMs encode vast biographical information, they exhibit task-specific instability under meaning-preserving transformations, with substantially larger degradation for occupational (concept-level) than biographical (entity-level) inference. Within the tasks and domains studied, their performance proves fragile under minor linguistic variations, underscoring a critical gap between surface-level knowledge retrieval and deeper inferential stability.

Antonio de Jesús García-Chávez, Everardo Gutiérrez-López, D. Flores et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.