Large language model (LLM)-based multi-agent systems show strong potential for supporting complex reasoning and decision-making in dialogue tasks. However, existing systems often rely on centralized coordination or uncontrolled inter-agent communication, which can limit scalability and increase communication overhead in multi-domain task-oriented dialogue environments. In this paper, we propose a decentralized multi-agent framework that combines role-based collaboration with communication-efficient coordination, enabling agents to operate effectively under constrained token budgets. The proposed approach aims to improve scalability, reduce communication cost, and enhance task success across diverse dialogue domains. Experimental results demonstrate that the proposed method outperforms single-agent and centralized coordination baselines, achieving a 15% improvement in task success rate and a 20% reduction in communication cost. These findings indicate that communication-efficient decentralized coordination significantly enhances both efficiency and robustness. Overall, the proposed architecture provides a practical and scalable solution for multi-domain task-oriented dialogue systems, particularly in resource-constrained environments.
Madhusundar Nelson· International journal of res...· 0 citations
Vision-language models have made remarkable improvements in multimodal understanding still, correctly mapping an image with its respective text is still problematic. A caption-based multimodal fusion scheme that fuses Bootstrapping Language-Image Pre-training Version 2(BLIP-2) and Contrastive Language–Image Pre-training(CLIP) to boost image-text alignment capabilities. This hypothesis lies on captions serving as semantic bridging for multimodal information. In the first stage, it utilizes BLIP-2 to generate a caption from the input image to extract semantic information at higher levels than pure visual information. Simultaneously, use CLIP to embed the image and text into a common space resulting in image vector and text vector. Subsequently, it encodes the generated caption through the CLIP text encoder to get semantic-aligned text vectors. Lastly, compute similarity scores between the image embeddings and text embeddings based on their semantic alignment. To increase representation power, this fusion block combines the semantic representation obtained from the caption generation process with the visual embeddings to allow a multimodal model to perceive both the visual and semantic information. This approach is evaluated on the Multi30K dataset, where image-text and text-image retrieval experiments were conducted. Conclusively, now it presents a scalable solution for improving multimodal alignment through the use of captions as a middle-level semantic representation.
Priyanka R, Madhusundar Nelson, D.Prabhu et al.· 2026 4th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.