Skip to content
Conference

Preventing Control Loss in Stochastic Environments: Recurrent Proximal Policy Optimization with Conditional Value at Risk

Sep 2026 · Automation, Control, and Information Technology · pp. 261-264 · 0 citations · 19 references

Abstract

Preventing transient control loss in risk-critical human-machine systems requires dynamic intervention strategies that account for unobservable user states. Traditional control frameworks relying on fully observable Markov decision processes are inadequate for this task, as they inherently suffer from perceptual aliasing when dealing with latent psychological variables. To address this computational challenge, this study introduces a risk-sensitive recurrent reinforcement learning framework that models the environment as a partially observable Markov decision process. By leveraging long short-term memory networks, the proposed control agent aggregates sequences of noisy behavioral indicators into comprehensive temporal trajectories, allowing for the accurate inference of hidden risk levels prior to critical failures. Furthermore, a composite reward-shaping mechanism is developed, replacing standard expected return maximization with a Conditional Value at Risk objective. This optimization strategy forces the policy to strictly penalize heavy-tailed risks and worst-case scenarios of behavioral escalation. Ultimately, the introduced computational model provides a robust algorithmic foundation for shifting from static risk classification to proactive, adaptive user support in highly stochastic environments.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.