In silico optimization of deep brain stimulation to enhance cognitive control: Improving performance and practicality with a continuous rolling arena
Abstract
Objective The use of Deep Brain Stimulation (DBS) on the ventral capsule/ventral striatum (VCVS) has therapeutic potential for patients with refractory psychiatric disorders, but clinical success is impeded by the need for a time-consuming and trial-and-error process when setting the parameters, this process relying on subjective self-reports. By monitoring objective behavioural markers it is possible to quickly assess the contact settings. In this study, we assess a direct closed-loop Multi-Armed Bandit (MAB) optimisation framework which is designed to quickly determine the best stimulation contacts by using raw reaction times (RT) obtained during a cognitive control task. Approach We leveraged empirical data demonstrating that VCVS DBS enhances cognitive control during the Multi-Source Interference Task (MSIT) in a site-specific manner. Using a synthetic patient simulation environment across 1,000 replicates, we benchmarked adaptive MAB algorithms under noisy, non-stationary conditions. Crucially, we eliminated intermediate state-space sensor models to evaluate raw RT directly, transitioned from discrete daily resets to an uninterrupted continuous optimization architecture, and implemented a rolling arena mechanism to scale contact selection under real-world hardware constraints. Main results Eliminating the intermediate sensor model prevented high-frequency noise amplification (where state variance was inflated by 51.7% in baseline and 148.0% in conflict states) and reduced contact ranking failure rates from 31.9% down to 11.9%. Operating within a continuous trial architecture preserved historical sample density, driving mean trial-level regret down steadily over 4,200 trials and enabling dynamic re-convergence across unannounced mid-session change-points. Additionally, a 4-contact sub-arena successfully scaled search efficiency across 8-contact arrays without sacrificing selection accuracy. Significance Direct MAB optimization within a continuous rolling arena provides a noise-resilient, hardware-compatible architecture for automated DBS programming. By bypassing latent state estimation and utilizing standard task-based behavioral metrics without specialized recording hardware or complex state-space modeling, this framework reduces search timelines to clinically feasible durations, establishing a scalable foundation for real-time, patient-tailored neuromodulation.