Jul 2026· International Journal of Food Science & Technology· Vol 61· 0 citations
TL;DR
It is suggested that off-the-shelf LLMs require prompt calibration and human oversight before UPF surveillance workflows, and low inter-run agreement indicated sensitivity to model updates, supporting prompt calibration and human oversight.
Abstract
This cross-sectional study compared three LLMs (Grok 4.1, Gemini 3, and ChatGPT 5.2) in classifying ultra-processed foods (UPF) using best-selling products from leading supermarket chains covering 53.2% of the national market. Of 3001 products, 2920 with complete ingredient information were included; two trained dietitians assigned NOVA groups as the reference standard. In the reference classification, 74.3% of products were UPF. Under the baseline prompt, all models underestimated UPF prevalence compared with the reference standard (p< 0.001). ChatGPT 5.2 yielded the highest binary UPF detection performance (accuracy: 69.01%; sensitivity: 59.01%; specificity: 98.00%; F1: 73.88%). Prompt sensitivity analyses revealed that a minimal prompt substantially outperformed the detailed baseline for most models (Gemini 3 F1: 94.20%; ChatGPT 5.2 F1: 92.62%). Low inter-run agreement (κ: 0.01–0.27) indicated sensitivity to model updates, supporting prompt calibration and human oversight. These findings suggest that off-the-shelf LLMs require prompt calibration and human oversight before UPF surveillance workflows.
Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.
Jiwon You, Hangsik Shin· npj Digital Medicine· 0 citations
Publicly available LLMs can pose potential safety risks to dental patients when used without guardrails and language-specific model training and hallucination reduction strategies, such as retrieval-augmented generation, are recommended.
Martyna Mysior, M. Mysior, Pamela Maslowski et al.· BMC Oral Health· 0 citations
Background/Objectives: The objective evaluation of regional dietary intake remains a core challenge in personalized health management due to complex plate presentations and a lack of culturally specific dataset benchmarks. Methods: This study introduces a confidence-aware hybrid vision–language framework engineered for traditional Turkish food recognition and structured nutritional assessment. Results: We curate a balanced dataset containing 14,711 verified images spanning 40 representative Turkish culinary classes to train and evaluate seven deep learning architectures. Among the visual models, EfficientNet V2-L achieved the highest standalone performance with an accuracy of 93.47%, 0.92 macro-precision, 0.92 macro-recall, and a 0.92 F1 score. To overcome visual ambiguity and automate content analysis, a confidence-aware routing strategy escalates uncertain predictions (τ<0.70) or user-rejected classifications to the Google Gemini 2.5 Flash multimodal large language model (MLLM). Conclusions: This hybrid paradigm yields a combined classification accuracy of 95.50% while validating portion weight estimations within a mean absolute error (MAE) of 18.42 g and total energy within 36.75 kcal. Fully realized as a cross-platform Flutter mobile application, the end-to-end pipeline demonstrates localized plate detection, adaptive portion analysis, and structured nutrient tracking, providing a scalable design for consumer-facing digital nutrition platforms.
Furkan Göz, Muhammad Jamil, A. Kavak et al.· Nutrients· 0 citations
Poor diet is now the leading cause of early death globally. In part, this is because our complex food supply chains are increasingly at risk of overprocessing, contamination, low nutrient content, and economically motivated fraud. Chemical testing can offer insights into these concerns, but testing methods are frequently impractical. Extra virgin olive oil (EVOO) is a premium food of high nutritional value, but because of its growing popularity and high price, it can be a target for mislabeling, substitution, dilution, and/or false claims of origin. Rapid and accurate testing methods for its characterization are therefore increasingly important. We used two direct forms of mass spectrometry (MS)-laser desorption/ionization (LDI) MS and direct analysis in real time (DART) MS-to obtain complex chemical signatures of edible oils. The data generated on a set of reference samples were then used to develop and train three independent machine learning (ML) models that assess key characteristics of a test oil. We also developed a proof-of-concept DART-MS/MS assay add-on for the quantification of bioactive phenols in EVOO. Our approach accurately predicts several attributes of an edible oil based on novel markers and intricate patterns within the acquired data. Further, pure reference standards and an isotopically-labeled internal standard allow accurate quantification of the constituent phenols. Because there is no chromatography, both the mass fingerprints and quantification can be performed in seconds-minutes. The method uses low (milliliter) volumes of sample and green solvents, and when combined with ML, it offers rapid data analysis and comprehensive result interpretation.
Nandhini Sokkalingam, Frances Chu, Joan Zou et al.· Journal of Mass Spectrometry· 0 citations
This study aimed to determine the prevalence of phosphorus (P)-based additives in processed foods and beverages in the Turkish market and evaluate how these components are declared on ingredient labels. Ingredient lists of 3,293 products across 16 food categories from eight major retail chains and one online market operating across Türkiye were systematically screened for phosphorus-based additives between January and April 2025. Data collected included food category, additive type (E number/name), total additive count, and declaration methods of phosphorus-based additives. P-based additives were identified in 58.3% of products. The highest prevalence was observed in cereal products (91.4%), ice creams (82.5%), and coffee and chocolate drinks (79.6%). Fourteen different P-based additives were detected, with lecithin (E 322, 39.9%), phosphate-containing modified starches (20.5%), and diphosphates (E 450, 18.8%) being the most common. Riboflavin-5'-phosphate was frequently found in snacks, while ammonium phosphatide was prominent in confectionery, cereal products, and ice creams. Declaration by name only (35.8%) was nearly twice as common as declaration by E number (18.2%). The prevalence of P-based additives in processed foods in Türkiye exceeds global averages. The widespread presence of highly bioavailable inorganic phosphates and 'hidden' sources such as modified starch and lecithin, combined with complex labelling practices, may pose risks, particularly for individuals with renal disease. Improved labelling policies are warranted to support public health.
Volkan Özkaya, E. Karabudak, Zehra Doğan· Food Additives and Contamina...· 0 citations
The question, “What are the best sources of protein?” is complicated by the significant hidden health and environmental costs associated with many popular options. Existing food rating systems, such as nutritional traffic lights and carbon footprint labels, are valuable but often unidimensional. This narrow focus can conceal critical trade-offs and mislead consumers by failing to present a holistic picture. This perspective paper introduces the BPI Score, a new, multidimensional rating system designed to provide a more comprehensive answer. The goal is not necessarily to create another front-of-package label, but to foster a new literacy that empowers consumers, policymakers, and organizations to make more informed decisions. The initial version (V1) of the BPI Score was used to evaluate 20 high-protein products across two primary categories: people and planet. The Planet Score assesses environmental impact using data on GHG emissions, water pollution, and resource use. The People Score considers health factors like sodium and saturated fat content alongside accessibility metrics such as affordability and availability. A clear disparity emerged between plant-based and animal-based proteins. Tofu emerged as the highest-scoring product, while cheddar cheese ranked last. The top quintile of products consisted solely of plant-based proteins, while the bottom quintile was composed entirely of animal-based proteins. While acknowledging the limitations of this first iteration, the score provides a robust foundation for a more nuanced conversation. Ultimately, beyond the spreadsheets and scores, lies a fundamental reckoning: our protein choice is a referendum on the compassion we are willing to show to the planet and to future generations.
C. MacDonald· Frontiers in Sustainable Foo...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.