Skip to content

Author

A. B. Banu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

ViT-FoodNA: An End-to-End Transformer-Based Multimodal Framework for Food Recognition and Nutrition Analysis

Dietary monitoring and health care will both require accurate food recognition and nutritional analysis. The current methods mainly involve convolutional networks that analyse images and word embedding that process ingredients as distinct processes with dish identification, portion size, and nutrition labelling as independent. This paper presents a multimodal transformer-based framework called as ViT-FoodNA that integrates these tasks into one end-to-end framework. It suggests using Vision Transformer (ViT), which is used to encode the visual attributes of food images, and Transformer based encoder to learn semantic relationships between inputs of ingredients. A cross-modal fusion transformer consists of visual and textual embedding to facilitate contextual interaction between the images of foods and their ingredients. Based on the result of the fused representation, the model collectively predicts the identity of the dish along with portion size and comprehensive nutritional breakdown using a shared transformer decoding strategy. In contrast to the previous convolutional neural network (CNN)-based embedding structures, the suggested architecture uses the self-attention mechanism to learn fine-grained intermodal dependencies, which enhances its resistance to visual ambiguity and missing ingredients in a list. Experiments indicate that ViT-FoodNA produces up to 12% higher Precision, Recall, F1-score, and 9.75% and 5% greater Top-1 and Top-5 accuracy and up to 32.7% and 30.4% lower MAE and RMSE than existing models. The reduction of PMAE is 6.5%, and the increase of mAP is 9.8%, which proves the great benefit of transformer-based multimodal fusion in the full analysis of diet.

E. Anitha, A. B. Banu · 0 citations
Open access Aug 2026

Biometrically Anchored and Psychological States Aware Deep Learning for Enhanced Personalized Food Recommendation Systems

The current Food Recommendation System (FRS) need new methods that combine psychological and biometric data about people to develop personalized food recommendations. The existing models use static user profiles and historical interaction data, which leads to their failure to capture the dynamic context-sensitive elements that drive user behaviour. To overcome these issues, the study proposed novel FRS, uses deep learning (DL) methods to deliver personalized food suggestions which depend on user characteristics, including their gender, mood and body weight. The architectural design uses body weight as a key metabolic reference point through which it establishes physically suitable recommendations that match the behavioural patterns associated with different gender and emotional states. The proposed method includes multiple essential steps, which start with dataset collection and preparation before proceeding to create dense vector representations, which reduce dimensionality and then use Convolutional Neural Networks (CNN) to extract food-related textual features, which lead to the discovery of essential food-related text patterns. And finally, create a multi-model feature fusion which combines user preferences with biometric constraints and food attributes and then generates Top-N recommendations through a Variational AutoEncoder (VAE). The framework introduces its novel aspect through a multi-modal latent representation system, which combines temporary emotional states with metabolic needs to create health-focused recommendations that adapt in real time. The VAE-based FRS (VAEFRS) system achieved optimization through its training process, which used a Conditional Tabular Generative Adversarial Network (CTGAN)-augmented training manifold to obtain full coverage of high-dimensional features while preventing overfitting issues. Experimental results show that the proposed model achieved a hit rate@10 of 0.8929, NDCG@10 of 0.6475, precision@10 (0.3223) and recall@10 (0.3533). The results demonstrate the effectiveness of the system in delivering pertinent and accurate FRs depending on their gender, physical characteristics, and present mood.

E. Anitha, A. B. Banu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.