Skip to content

Author

jaeha shin

6 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#reinforcement learning Open access Sep 2026

Reinforcement learning optimization of a ⁶Li magneto-optical trap: DDPG versus SAC

We report an experimental implementation of deep reinforcement learning (RL) for optimizing the loading stage of a ⁶Li magneto-optical trap (MOT) in a high-dimensional continuous control space. An off-policy actor-critic agent observes in-situ fluorescence images at 111 ms intervals and updates seven experimental parameters over 18 decision steps. Learning is performed directly on the apparatus using a sparse terminal reward 𝑅 = 𝑁 × 𝐴 extracted from an absorption image, where 𝑁 is the atom number and 𝐴 is the peak optical depth. We benchmark Deep Deterministic Policy Gradient (DDPG) and Soft Actor-Critic (SAC) across three independent 300-episode training runs under identical alignment. Both algorithms surpass the human-optimized (HO) baseline, but the central finding concerns reliability rather than peak performance. In deterministic replay validation, the three SAC waveforms show reward coefficients of variation of 9–14%, and the minimum reward of each waveform exceeds the HO mean reward. The DDPG waveforms show up to three times the dispersion, and even the best-performing DDPG run contains a catastrophic near-zero shot. The SAC solutions are nearly time-independent and consistently converge to a blue-detuned repump laser with an elevated field gradient, reminiscent of compressed-MOT operating conditions. Our results show that off-policy actor-critic RL can autonomously optimize high-dimensional laser-cooling sequences in an operating ultracold-atom apparatus, and that SAC has a reproducibility advantage over DDPG in a real noisy environment.

Deok-Young Lee, jaeha shin, Jae-yoon Choi · 0 citations
#reinforcement learning Open access Sep 2026

Reinforcement learning optimization of a ⁶Li magneto-optical trap: DDPG versus SAC

We report an experimental implementation of deep reinforcement learning (RL) for optimizing the loading stage of a ⁶Li magneto-optical trap (MOT) in a high-dimensional continuous control space. An off-policy actor-critic agent observes in-situ fluorescence images at 111 ms intervals and updates seven experimental parameters over 18 decision steps. Learning is performed directly on the apparatus using a sparse terminal reward 𝑅 = 𝑁 × 𝐴 extracted from an absorption image, where 𝑁 is the atom number and 𝐴 is the peak optical depth. We benchmark Deep Deterministic Policy Gradient (DDPG) and Soft Actor-Critic (SAC) across three independent 300-episode training runs under identical alignment. Both algorithms surpass the human-optimized (HO) baseline, but the central finding concerns reliability rather than peak performance. In deterministic replay validation, the three SAC waveforms show reward coefficients of variation of 9–14%, and the minimum reward of each waveform exceeds the HO mean reward. The DDPG waveforms show up to three times the dispersion, and even the best-performing DDPG run contains a catastrophic near-zero shot. The SAC solutions are nearly time-independent and consistently converge to a blue-detuned repump laser with an elevated field gradient, reminiscent of compressed-MOT operating conditions. Our results show that off-policy actor-critic RL can autonomously optimize high-dimensional laser-cooling sequences in an operating ultracold-atom apparatus, and that SAC has a reproducibility advantage over DDPG in a real noisy environment.

Deok-Young Lee, jaeha shin, Jae-yoon Choi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.