Skip to content
Open access

Can Machines Detect Ultra-Processed Foods? A Head-to-Head Evaluation of Large Language Models Using NOVA Classification

Jul 2026 · International Journal of Food Science & Technology · Vol 61 · 0 citations

TL;DR

It is suggested that off-the-shelf LLMs require prompt calibration and human oversight before UPF surveillance workflows, and low inter-run agreement indicated sensitivity to model updates, supporting prompt calibration and human oversight.

Abstract

This cross-sectional study compared three LLMs (Grok 4.1, Gemini 3, and ChatGPT 5.2) in classifying ultra-processed foods (UPF) using best-selling products from leading supermarket chains covering 53.2% of the national market. Of 3001 products, 2920 with complete ingredient information were included; two trained dietitians assigned NOVA groups as the reference standard. In the reference classification, 74.3% of products were UPF. Under the baseline prompt, all models underestimated UPF prevalence compared with the reference standard (p< 0.001). ChatGPT 5.2 yielded the highest binary UPF detection performance (accuracy: 69.01%; sensitivity: 59.01%; specificity: 98.00%; F1: 73.88%). Prompt sensitivity analyses revealed that a minimal prompt substantially outperformed the detailed baseline for most models (Gemini 3 F1: 94.20%; ChatGPT 5.2 F1: 92.62%). Low inter-run agreement (κ: 0.01–0.27) indicated sensitivity to model updates, supporting prompt calibration and human oversight. These findings suggest that off-the-shelf LLMs require prompt calibration and human oversight before UPF surveillance workflows.

Read PDF

Similar papers

Open access Jul 2026

Large language models for interpretation of health checkup results

Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.

Jiwon You, Hangsik Shin · 0 citations
Open access Jul 2026

A cross-sectional evaluation of Large Language Model answers to dental questions.

Publicly available LLMs can pose potential safety risks to dental patients when used without guardrails and language-specific model training and hallucination reduction strategies, such as retrieval-augmented generation, are recommended.

Martyna Mysior, M. Mysior, Pamela Maslowski et al. · 0 citations
Open access Jul 2026

A Confidence-Aware Hybrid Vision–Language Framework for Food Recognition and Nutritional Monitoring

Background/Objectives: The objective evaluation of regional dietary intake remains a core challenge in personalized health management due to complex plate presentations and a lack of culturally specific dataset benchmarks. Methods: This study introduces a confidence-aware hybrid vision–language framework engineered for traditional Turkish food recognition and structured nutritional assessment. Results: We curate a balanced dataset containing 14,711 verified images spanning 40 representative Turkish culinary classes to train and evaluate seven deep learning architectures. Among the visual models, EfficientNet V2-L achieved the highest standalone performance with an accuracy of 93.47%, 0.92 macro-precision, 0.92 macro-recall, and a 0.92 F1 score. To overcome visual ambiguity and automate content analysis, a confidence-aware routing strategy escalates uncertain predictions (τ<0.70) or user-rejected classifications to the Google Gemini 2.5 Flash multimodal large language model (MLLM). Conclusions: This hybrid paradigm yields a combined classification accuracy of 95.50% while validating portion weight estimations within a mean absolute error (MAE) of 18.42 g and total energy within 36.75 kcal. Fully realized as a cross-platform Flutter mobile application, the end-to-end pipeline demonstrates localized plate detection, adaptive portion analysis, and structured nutrient tracking, providing a scalable design for consumer-facing digital nutrition platforms.

Furkan Göz, Muhammad Jamil, A. Kavak et al. · 0 citations
Open access Aug 2026

A Platform for High-Throughput Chemical Analysis of Foods: Characterization of Extra Virgin Olive by Direct Mass Spectrometry and Machine Learning.

Poor diet is now the leading cause of early death globally. In part, this is because our complex food supply chains are increasingly at risk of overprocessing, contamination, low nutrient content, and economically motivated fraud. Chemical testing can offer insights into these concerns, but testing methods are frequently impractical. Extra virgin olive oil (EVOO) is a premium food of high nutritional value, but because of its growing popularity and high price, it can be a target for mislabeling, substitution, dilution, and/or false claims of origin. Rapid and accurate testing methods for its characterization are therefore increasingly important. We used two direct forms of mass spectrometry (MS)-laser desorption/ionization (LDI) MS and direct analysis in real time (DART) MS-to obtain complex chemical signatures of edible oils. The data generated on a set of reference samples were then used to develop and train three independent machine learning (ML) models that assess key characteristics of a test oil. We also developed a proof-of-concept DART-MS/MS assay add-on for the quantification of bioactive phenols in EVOO. Our approach accurately predicts several attributes of an edible oil based on novel markers and intricate patterns within the acquired data. Further, pure reference standards and an isotopically-labeled internal standard allow accurate quantification of the constituent phenols. Because there is no chromatography, both the mass fingerprints and quantification can be performed in seconds-minutes. The method uses low (milliliter) volumes of sample and green solvents, and when combined with ML, it offers rapid data analysis and comprehensive result interpretation.

Nandhini Sokkalingam, Frances Chu, Joan Zou et al. · 0 citations
Review Aug 2026

A survey of selected packaged food products in Türkiye for listed ingredients containing phosphorus-based food additives.

This study aimed to determine the prevalence of phosphorus (P)-based additives in processed foods and beverages in the Turkish market and evaluate how these components are declared on ingredient labels. Ingredient lists of 3,293 products across 16 food categories from eight major retail chains and one online market operating across Türkiye were systematically screened for phosphorus-based additives between January and April 2025. Data collected included food category, additive type (E number/name), total additive count, and declaration methods of phosphorus-based additives. P-based additives were identified in 58.3% of products. The highest prevalence was observed in cereal products (91.4%), ice creams (82.5%), and coffee and chocolate drinks (79.6%). Fourteen different P-based additives were detected, with lecithin (E 322, 39.9%), phosphate-containing modified starches (20.5%), and diphosphates (E 450, 18.8%) being the most common. Riboflavin-5'-phosphate was frequently found in snacks, while ammonium phosphatide was prominent in confectionery, cereal products, and ice creams. Declaration by name only (35.8%) was nearly twice as common as declaration by E number (18.2%). The prevalence of P-based additives in processed foods in Türkiye exceeds global averages. The widespread presence of highly bioavailable inorganic phosphates and 'hidden' sources such as modified starch and lecithin, combined with complex labelling practices, may pose risks, particularly for individuals with renal disease. Improved labelling policies are warranted to support public health.

Volkan Özkaya, E. Karabudak, Zehra Doğan · 0 citations
Open access Jul 2026

What are the best sources of protein? Introducing the BPI Score

The question, “What are the best sources of protein?” is complicated by the significant hidden health and environmental costs associated with many popular options. Existing food rating systems, such as nutritional traffic lights and carbon footprint labels, are valuable but often unidimensional. This narrow focus can conceal critical trade-offs and mislead consumers by failing to present a holistic picture. This perspective paper introduces the BPI Score, a new, multidimensional rating system designed to provide a more comprehensive answer. The goal is not necessarily to create another front-of-package label, but to foster a new literacy that empowers consumers, policymakers, and organizations to make more informed decisions. The initial version (V1) of the BPI Score was used to evaluate 20 high-protein products across two primary categories: people and planet. The Planet Score assesses environmental impact using data on GHG emissions, water pollution, and resource use. The People Score considers health factors like sodium and saturated fat content alongside accessibility metrics such as affordability and availability. A clear disparity emerged between plant-based and animal-based proteins. Tofu emerged as the highest-scoring product, while cheddar cheese ranked last. The top quintile of products consisted solely of plant-based proteins, while the bottom quintile was composed entirely of animal-based proteins. While acknowledging the limitations of this first iteration, the score provides a robust foundation for a more nuanced conversation. Ultimately, beyond the spreadsheets and scores, lies a fundamental reckoning: our protein choice is a referendum on the compassion we are willing to show to the planet and to future generations.

C. MacDonald · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.