Aug 2026· Academic Journal of Research and Scientific Publishing· Vol 8, pp. 5-42· 0 citations
TL;DR
The study concludes that future LLM systems are likely to become more adaptive, verifiable, self-evaluating, and self-correcting, and recommends improving judge reliability, error traceability, and human oversight in high-risk applications, while exploring judge-governed execution as a future extension for agentic and robotic systems.
Abstract
This study aims to develop an integrative computational-analytical framework for explaining the functional evolution of large language models through four interrelated dimensions: Reasoning, Retrieval and Knowledge Grounding, Alignment, and Judging, collectively represented by the RRAJ framework. The study adopts a systematic analytical literature review of research published between 2020 and 2026, drawing on Scopus, Web of Science, and IEEE Xplore, and applying comparative and thematic analysis to trace major technical transformations and functional relationships across the four dimensions. The findings indicate that the evolution of large language models is no longer driven primarily by model scaling, but increasingly by system scaling, in which knowledge, computation, and control are distributed across model parameters, contextual information, external knowledge sources, inference-time computation, and evaluation mechanisms. The analysis further reveals a shift from isolated capabilities toward functional convergence among RRAJ components, with LLM-as-a-Judge emerging as an important transition that enables evaluation to support diagnosis, feedback, and correction. The study concludes that future LLM systems are likely to become more adaptive, verifiable, self-evaluating, and self-correcting. It further recommends improving judge reliability, error traceability, and human oversight in high-risk applications, while exploring judge-governed execution as a future extension for agentic and robotic systems.
Large language models represent a qualitatively new development within generative artificial intelligence and have recently begun to appear in circular economy research, yet no review has pulled together what is actually known about how and where these tools are being used in this field. This study aims to map current applications of large language models and generative artificial intelligence in circular economy contexts, examine how these tools are being deployed methodologically, and identify what the evidence says about their benefits, limitations, and risks. Following PRISMA-ScR guidance, a scoping review was conducted in Scopus in February 2026, returning 127 records, which were screened down to 30 peer-reviewed journal articles published between 2024 and 2026 and analyzed through a matrix-based coding approach. The findings show that large language models are used less as standalone decision-making tools and more as ways of organizing and mobilizing knowledge for practitioners and researchers, with roles ranging from design assistants and knowledge integrators to analytical components in larger modelling workflows. Most applications are found in industrial symbiosis, circular supply chains, sustainable product and materials innovation, and the built environment, while newer work is beginning to emerge in waste sorting, policy analysis, and skills mapping. The studies consistently show that the most reliable results come when large language models are combined with domain-specific knowledge bases and expert validation, rather than used autonomously. Across these application areas, several governance challenges remain unaddressed, including the absence of domain-specific validation standards, insufficient data infrastructure for circular economy contexts, and the largely unaccounted environmental footprint of generative AI hardware itself. Based on these findings, future research should focus on comparative evaluations and better integration with established tools such as life cycle assessment. Policymakers need to invest in domain-specific data infrastructure and clear governance frameworks, and practitioners should treat human oversight as a necessary part of any workflow that uses these tools.
It is argued that understanding AI knowledge creation is essential for bridging traditional human KM with the emerging discipline of AI Knowledge Management, and for designing governance structures that account for the inherent incompleteness and inconsistency of LLM knowledge.
T. Nguyen· European Conference on Knowl...· 0 citations
The rapid evolution of Large Language Models (LLMs) has brought unprecedented capabilities across reasoning, coding, and multimodal tasks. However, as performance scales, their opaque ''black-box'' nature raises a critical challenge: How can we trace the origins of emergent intelligence, and more importantly, how can we leverage these internal mechanisms to guide model optimization? This tutorial provides a comprehensive, end-to-end view of LLM interpretability, transitioning from microscopic neural analysis to macroscopic application and deployment. It is systematically organized into five core sections: i) Unlocking the Black Box: We begin with the evolution of LLM interpretability and highlight recent breakthroughs from leading research teams. ii) Methodology: We present a rigorous overview of foundational theories (e.g., mathematical framework for transformer, biological mechanisms in LLMs) and essential methods (e.g., path patching, logit lens, and neuron description). iii) Anatomy of LLMs: Using advanced techniques to decode internal semantic features, neural circuits, and complex behaviors, we interpret how models perform reasoning, factual recall, and in-context learning. iv) Applications: We show how to transfer interpretability insights into actionable improvements across the LLM pipeline, including interpretability-guided data synthesis (data value scoring, corpus filtering, and activation-based data diagnosis). We also present Pinpoint Training and Steering for precise capability gains, and Pinpoint Quantization for extreme low-bit compression with minimal capability loss. v) Advanced Topics: We conclude by exploring how these interpretability paradigms scale and inspire the design of frontier architectures, agentic systems, and thinking models. In this tutorial, researchers and engineers will gain the theoretical frameworks and practical engineering toolkits needed to understand, steer, and efficiently deploy LLMs in real-world production environments.
Wei Zhang, Zhengfu He, Lucia Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
This paper introduces a comprehensive multi-criteria evaluation methodology designed to assess the capabilities of these advanced computational architectures in handling complex language tasks, focusing on three foundational dimensions: hierarchical reasoning, self-correction mechanisms, and factual consistency.
S. Yam· Journal of innovative resear...· 0 citations
Despite the expansion of research on leadership behaviors (LBs) over the past decades, this field faces challenges related to construct redundancy, which impedes theoretical integration and empirical accumulation. To address this issue, we conducted two complementary studies to systematically examine a six-category taxonomy of LBs (i.e., task-oriented, relational-oriented, change-oriented, value-based and moral, passive, and destructive LBs, collectively referred to as the TRCVPD taxonomy), which maps diverse LBs onto a unified framework. In Study 1, we compiled a database comprising measurement items from 48 LBs scales and leveraged large language models and subject-matter experts to assess whether measurement items of previously unclassified LBs semantically corresponded to the theoretical definitions of the TRCVPD taxonomy of LBs. Specifically, we were able to semantically match 41 of the 48 LBs to the TRCVPD taxonomy. Based on the findings of Study 1, Study 2 conducted a comprehensive second-order meta-analysis, synthesizing 109 meta-analyses based on more than 6 million participants. In particular, we evaluated the nomological network of the six LBs by meta-analytically estimating their relations with antecedents and outcomes. In addition, we used dominance analyses to further compare the relative importance of each LB category in explaining leadership outcomes. Together, these two studies present integrative findings that offer an overarching theoretical framework and empirical foundation for understanding the taxonomy and nomological structure of LBs. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Yucheng Zhang, Shanshan Zhang, Zijun Ke et al.· Journal of Applied Psycholog...· 0 citations
This study proposes a conceptual modeling approach based on Multi-Criteria Evaluation (MCE) for the analysis of complex systems with interacting heterogeneous factors, formulated in semi-quantitative terms, in which determining an optimal balance is a central modeling problem. The approach integrates structural analysis, influence modeling, and utility-based evaluation into a unified methodology that enables both interpretation and quantitative assessment of system behavior.
The proposed approach is applied to a case study of term formation in IT terminology, where linguistic and cognitive parameters are treated as system variables. Two competing scenarios are modeled: structure-oriented formation, driven by morphological regularities, and economy-driven formation, governed by cognitive efficiency principles. By mapping these scenarios into a multidimensional utility space, the model enables the identification of trade-offs and optimal configurations.
The results demonstrate that, in the context of IT terminology, the economy-driven principle dominates the overall system evaluation, indicating that communicative efficiency and compression play a decisive role in shaping term formation under contemporary conditions. At the same time, structural factors contribute to system stability through balancing feedback mechanisms.
V. Dmytruk, M. Voloshyn, O. Senkovych et al.· Mathematical Modeling and Co...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.