Skip to content
Open access

ViT-FoodNA: An End-to-End Transformer-Based Multimodal Framework for Food Recognition and Nutrition Analysis

2026 · ITEGAM- Journal of Engineering and Technology for Industrial Applications (ITEGAM-JETIA) · 0 citations

Abstract

Dietary monitoring and health care will both require accurate food recognition and nutritional analysis. The current methods mainly involve convolutional networks that analyse images and word embedding that process ingredients as distinct processes with dish identification, portion size, and nutrition labelling as independent. This paper presents a multimodal transformer-based framework called as ViT-FoodNA that integrates these tasks into one end-to-end framework. It suggests using Vision Transformer (ViT), which is used to encode the visual attributes of food images, and Transformer based encoder to learn semantic relationships between inputs of ingredients. A cross-modal fusion transformer consists of visual and textual embedding to facilitate contextual interaction between the images of foods and their ingredients. Based on the result of the fused representation, the model collectively predicts the identity of the dish along with portion size and comprehensive nutritional breakdown using a shared transformer decoding strategy. In contrast to the previous convolutional neural network (CNN)-based embedding structures, the suggested architecture uses the self-attention mechanism to learn fine-grained intermodal dependencies, which enhances its resistance to visual ambiguity and missing ingredients in a list. Experiments indicate that ViT-FoodNA produces up to 12% higher Precision, Recall, F1-score, and 9.75% and 5% greater Top-1 and Top-5 accuracy and up to 32.7% and 30.4% lower MAE and RMSE than existing models. The reduction of PMAE is 6.5%, and the increase of mAP is 9.8%, which proves the great benefit of transformer-based multimodal fusion in the full analysis of diet.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.