Aug 2026· International journal of information and communication technology trends· 0 citations
TL;DR
The research details a comprehensive methodological framework, formalizing the probabilistic decision-making and critique generation processes and indicates that integrating reflective cognition paradigms with modular toolsets is essential for deploying autonomous language agents in high-stakes, real-world applications.
Abstract
but struggle when confronted with multi-step, complex task planning that requires interaction with external environments. This paper investigates the architecture, implementation, and efficacy of tool-augmented language agents enhanced with iterative self-critique mechanisms. By integrating external application programming interfaces, structured databases, and computational engines, these agents transcend isolated text generation, evolving into active systems capable of executing concrete actions. However, naive tool utilization often results in cascading errors during prolonged execution trajectories. To mitigate this, we introduce an iterative self-critique framework where the agent continuously evaluates its own outputs, identifies logical fallacies or execution failures, and dynamically recalibrates its plan. This research details a comprehensive methodological framework, formalizing the probabilistic decision-making and critique generation processes. Empirical evaluations across simulated complex environments demonstrate that the proposed architecture significantly improves task success rates, minimizes superfluous tool invocations, and enhances error recovery. The findings indicate that integrating reflective cognition paradigms with modular toolsets is essential for deploying autonomous language agents in high-stakes, real-world applications.
Extensive experiments show that ExpG brings consistent improvements across the tool selection, tool calling, and response generation tasks, enabling smaller agents to outperform larger ones that do not use ExpG, suggesting a promising path toward more robust tool use.
Large Language Models (LLMs) have enabled the emergence of autonomous AI agents capable of
reasoning, planning, tool use, and iterative decision-making. Despite rapid development, the field
remains architecturally fragmented, with limited conceptual clarity regarding memory
integration, planning mechanisms, and operational reliability.
This study presents a systematic review and critical synthesis of LLM-based autonomous agents,
focusing on architectural paradigms, memory models, planning strategies, and real-world
deployment constraints. Using a structured review approach, this study examines existing LLM
based agent systems across key design components to uncover common patterns, differences in
implementation, and recurring structural weaknesses.
The review reveals persistent and structurally significant challenges across all four dimensions:
long-horizon reasoning stability degrades as task length increases; memory consistency is
undermined by retrieval noise, embedding drift, and summarisation errors; tool alignment failures
propagate errors across modular pipelines; and evaluation standardisation remains insufficient
to support reliable cross-paper comparison. A consistent cross-paradigm finding emerges:
autonomy and reliability trade off systematically as agent complexity increases, with current
systems achieving capability gains through heuristic design rather than principled theoretical
foundations. Based on this synthesis, the review proposes a consolidated analytical framework
that maps common structural elements and trade-offs across reviewed systems, and outlines a
research agenda directed toward formalised agent architectures, memory consistency guarantees,
verified planning algorithms, standardised reliability metrics, and benchmark frameworks
adequate for long-horizon, real-world evaluation conditions.
Unknown authors· International Journal of Com...· 0 citations
MUSE is presented, an interactive meta-agent that enhances user understanding and control of agentic data science systems by dynamically restructuring low-level execution traces into multiple semantic levels that support navigation from high-level overviews to low-level implementation details.
Wei-Hao Chen, Weixi Tong, Yuan Tian et al.· 0 citations
This research introduces a tiered multi-agent architecture grounded in Human-Centered eXplainable AI principles that contributes an adaptable and generalizable framework and foundational artifacts for trustworthy AI teammates.
Jie Tao, Li-Na Zhou· Information Systems Frontier...· 0 citations
A unified, taxonomy-driven, and deployment-oriented survey of agentic AI systems, synthesizing recent advances through a modular reference architecture and a four-dimensional taxonomy that characterizes agents along the axes of autonomy, tool use, collaboration, and safety–governance is presented.
Sparsh Bajoria, Shreyanshu Ranjan, Adhitya M et al.· Cognitive Computation· 0 citations
Multi-agent systems built on large language models (LLMs) are increasingly deployed for complex tasks requiring autonomous planning, tool use, and inter-agent coordination. However, the non-deterministic nature of LLM outputs and the emergent behavior arising from agent interactions render traditional test oracles ineffective, creating a critical gap in quality assurance for agentic AI. This work introduces MORPHAGENT, a framework designed to address the oracle problem in multi-agent LLM systems through trace-based behavioral analysis. Our contributions are threefold: (1) goal-preservation relations that verify consistent goal achievement under input perturbations, (2) coordination-consistency relations that validate inter-agent delegation and communication patterns under agent substitution and reordering, and (3) tool-use integrity relations that ensure semantic equivalence of tool invocation sequences under prompt paraphrasing. MorphAgent instruments agent execution to capture structured traces comprising planning steps, tool calls, message exchanges, and final outputs, then systematically applies metamorphic transformations and checks behavioral invariants without requiring ground-truth oracles. We evaluate the framework on four multi-agent benchmarks spanning code generation, research synthesis, customer service, and data analysis tasks, encompassing 2,840 source-followup execution pairs across three LLM backends. Results show that MORPHAGENT detects 82.0% of seeded behavioral faults, including 90.3% of coordination failures and 81.7% of goal-deviation faults, while maintaining a false positive rate of 6.1%. The framework uncovers 14 previously unreported behavioral anomalies in established multi-agent frameworks, demonstrating its practical utility for assuring agentic AI reliability. These results suggest that trace-based metamorphic testing can serve as a practical foundation for reliable validation of emerging agentic AI systems.
Gopalakrishnan Marimuthu· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.