2026· Computers, Materials & Continua· 0 citations· 26 references
TL;DR
Empirical support for domain-specialized structured reasoning and process-level evaluation as foundations for future RL-based C&D capability development in LLMs is provided.
Abstract
: Large Language Models (LLMs) currently lack the robust command and decision-making (C&D) capabilities essential for the command and control domain. To address this critical gap, this paper proposes an emergence mechanism that integrates a domain-specialized Chain of Thought (CoT) framework with a Process Reward Model (PRM)-inspired evaluation and inference-time optimization paradigm. We construct a novel Chain of Command and Decision (CoCD) framework, a C2-specific CoT structure with contextual persistence, knowledge accumulation, and a human-in-the-loop feedback loop, and define a four-dimensional PRM-inspired evaluation framework for process-level assessment of C&D reasoning. Experimental evaluations on 40 C&D scenarios of varying complexity demonstrate that the CoCD framework significantly outperforms direct prompting (Mann–Whitney U = 1314, p < 0.0001, Cohen’s d = 1.340) and Standard-CoT ( p = 0.005, d = 0.606) in composite performance. PRM-guided Best-of-N selection further improves performance by 5.8% over single-sample CoCD ( p < 0.001, d = 0.855), providing direct empirical evidence for the utility of process-aware reward signals at inference time. CoCD’s structural advantage is greatest in high-uncertainty, structurally ambiguous scenarios (Level 3 gap: + 0.925 points), revealing a complexity-type effect that informs the deployment scope of structured CoT frameworks. These findings provide empirical support for domain-specialized structured reasoning and process-level evaluation as foundations for future RL-based C&D capability development in LLMs.
A novel RAIE taxonomy along four scaling dimensions is proposed, which optimizes the entire thought process through search algorithms and self-verification, and introduces a task-oriented guideline for choosing the best TTS strategy.
Jia-Yu An, Zheng Chen, Yongcheng Jing et al.· 0 citations
This work proposes a novel neuro-symbolic fast-slow thinking (NeSyFS) framework for LLM agent, addressing the challenges introduced by partial observability in a unified approach, and uses a knowledge graph to represent the belief state.
PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents that achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines is introduced.
This work presents a theoretical framework that reveals how reasoning steps can amplify error through three failure modes: incorrect sub-task decomposition, incorrect sub-task solving, and incorrect final answer summarization, and introduces structured interventions that adapt CoT generation according to the identified failure types.
Haibo Jin, Peiyan Zhang, Man Luo et al.· Neural Information Processin...· 1 citation
This work introduces the Language Model Council (LMC), a collaborative framework that combines the expertise of multiple specialized AI agents to evaluate a user query from different perspectives and outperforms traditional single-model systems by improving response quality, reducing hallucinations, and increasing user trust through enhanced explainability.
D. M, Shwetha Kr, G. Divya et al.· International Research Journ...· 0 citations