Objectives: To develop a framework that integrates pedagogical structure with uncertainty-aware decision-making for personalized learning in smart educational environments, addressing the limitations of current deep reinforcement learning approaches that treat curricula as unstructured sequences. Method: The study formalizes the learning domain as a concept lattice—an order-theoretic structure derived from formal concept analysis that encodes prerequisite relationships. Within this structured state space, a Bayesian reinforcement learning agent using Thompson sampling maintains joint posterior distributions over the learner's latent knowledge state and the uncertain reward associated with each instructional action. The framework was evaluated on the ASSISTments 2012-2013 dataset (4,317 problems, 112 knowledge components, 334,416 interactions) and Eedi (98 concepts, 7,547 interactions)—and validated against four baseline methods: Standard Thompson Sampling, Graph-Constrained RL, Bayesian RL, and Static Policy. Findings: The proposed Structured Thompson Sampling (STS) framework achieved a 15.2% improvement in average skill gain over standard Thompson sampling on ASSISTments and a 14.8% improvement on Eedi, demonstrating consistent performance across datasets. The system demonstrated faster convergence with approximately 32% fewer training interactions. The system outputs well-calibrated uncertainty estimates with an expected calibration error of 0.036, supporting interpretable decision-making for educators. The Pedagogical Coherence Score of 0.96 confirms that STS respects prerequisite relationships, while ablation studies revealed that both the lattice structure and Bayesian optimization contribute significantly to performance. Novelty: This work presents the first integration of formal concept analysis with Bayesian reinforcement learning for pedagogical sequencing, providing a mathematically rigorous foundation for personalized learning that combines structural validity with quantifiable confidence estimates. The framework bridges the critical gap between pedagogical coherence and uncertainty-aware decision-making in adaptive educational systems.
Keywords: Bayesian Reinforcement Learning, Personalized Learning, Concept Lattice, Thompson Sampling, Pedagogical Sequencing, Uncertainty Quantification, Smart Learning Environments, Adaptive Educational Systems
S. Ahamed, A. R. Mohamed Shanavas· Indian Journal of Science an...· 0 citations
The rapid advancement of technology like Internet of Things (IoT) and Cloud computing (CC)based heterogeneous environment required dynamic resource management system. Thecomplexity of IoT-Cloud is increasing due to abundance of dynamic data dissemination thatcreate poor performance like high energy consumption (EC), workload imbalancing, pooradaptability, and fail to handle SLA violations. The two primary contributions of the proposedwork are the Optimized Priority-Aware Hierarchical Multi-Agent Deep Q Network (OPHMDQN)and the Adaptive Multi-Objective Dung Beetle Optimization Algorithm (ADBOA). Themulti-agent strategies increase the scalability and adaptability of resource managementthrough local and global hierarchies. The approach integrates information-based decisionmakingand priority-aware allocation while accounting for SLA requirements, systemconstraints, and job complexity to optimise resource generation, utilisation, allocation, andtask scheduling. In comparison to existing optimization and Reinforcement Learning (RL)techniques, experimental results show that the proposed OPHM-DQN-ADBOA frameworkconsistently reduces EC (up to 30 % lower), execution delay, and SLA violations whileimproving resource utilisation and LB. The ADBOA enhances the proposed model throughoptimal multi-objective training, reducing EC, SLA violation, and cost while improvingresource utilization and scheduling efficiency. As a results, the model achieves high scalabilityand adaptability in heterogeneous IoT-Cloud resource management.
A. Ali, A. R. Mohamed Shanavas· THE SCIENTIFIC TEMPER· 0 citations
Mobile Ad-hoc Network (MANET) are highly dynamic and infrastructure-less wireless network in which frequent topology changes, node mobility, packet collision, and energy constraints significantly affect routing performance and network reliability. This research suggests a Deep Reinforcement Learning (DRL)-based Optimized Multi-Path Relay Node Selection method for dependable and energy-efficient MANET routing. The suggested method takes into account important network metrics such as residual energy, node mobility, link stability, congestion level, and packet collision probability in order to intelligently choose the best relay nodes and different routing options using a Deep Q-Network (DQN)-based learning model. In comparison to traditional MANET routing protocols, simulation results show that the suggested DRL-based relay node selection technique greatly improves network lifetime, Packet Delivery Ratio (PDR), throughput and routing stability while lowering packet collision, end-to-end delay and energy consumption.
A. A. Samathu, G. Ravi, A. R. Mohamed Shanavas· THE SCIENTIFIC TEMPER· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.