Skip to content

Author

Stefan Zlatinov

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#reinforcement learning Open access Sep 2026

RoboSuite Grasping

This repository contains the source code accompanying the paper “Noise Scheduling and Out-of-Distribution Generalisation in Deep Reinforcement Learning for 6-DOF Robotic Grasping.” The work evaluates three off-policy deep reinforcement learning configurations for a 6-DOF robotic grasping task using a Universal Robots UR5e in the RoboSuite simulation environment: Soft Actor-Critic (SAC) with automatic entropy tuning, Twin Delayed Deep Deterministic Policy Gradient (TD3) with piecewise linear exploration-noise decay, and TD3 with constant Gaussian noise. The repository includes the implementation and experimental setup used to compare training efficiency, in-distribution grasping performance, and out-of-distribution generalisation across the three configurations. The study focuses particularly on the effect of exploration-noise scheduling on policy robustness outside the training distribution.

Stefan Zlatinov, Ilija Mizhimakoski · 0 citations
#reinforcement learning Open access Sep 2026

RoboSuite Grasping

This repository contains the source code accompanying the paper “Noise Scheduling and Out-of-Distribution Generalisation in Deep Reinforcement Learning for 6-DOF Robotic Grasping.” The work evaluates three off-policy deep reinforcement learning configurations for a 6-DOF robotic grasping task using a Universal Robots UR5e in the RoboSuite simulation environment: Soft Actor-Critic (SAC) with automatic entropy tuning, Twin Delayed Deep Deterministic Policy Gradient (TD3) with piecewise linear exploration-noise decay, and TD3 with constant Gaussian noise. The repository includes the implementation and experimental setup used to compare training efficiency, in-distribution grasping performance, and out-of-distribution generalisation across the three configurations. The study focuses particularly on the effect of exploration-noise scheduling on policy robustness outside the training distribution.

Stefan Zlatinov, Ilija Mizhimakoski · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.