Skip to content

Author

Zhennan Fan

We have 2 of 2 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning in Long-Horizon AntMaze Navigation

Long-horizon navigation requires a high-level policy to select locally reachable subgoals, yet a scalar task reward provides little information about how an unsuitable proposal should be changed. We introduce Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning (SCS-HRL), a two-level method in which a topology- and clearance-aware programmatic supervisor evaluates each proposed subgoal and returns both a scalar score and a continuous target in the same subgoal space. The score trains the high-level critic, and the target enters a masked regression term for the high-level actor. Primitive actions are always conditioned on the actor’s subgoal; the supervisor is inactive during learned-policy evaluation. In AntMaze, using 6000 training episodes, five seeds, and 100 deterministic evaluation episodes per seed, SCS-HRL attained an 88.4±7.8% final success rate (mean ± sample standard deviation; 95% Student-t confidence interval [78.7%,98.1%]). The matched scalar-only condition and HIRO attained 0% rates. Applying the same route rule directly to the SCS-HRL low-level controllers yielded 82.2±9.9% success; the paired difference favored the learned high-level policy by 6.2 percentage points (95% confidence interval [2.1,10.3], p=0.013). Across three matched seeds, nonzero corrective weights of 0.5, 1.0, and 2.0 remained stable, whereas 0.25 was seed-sensitive. Term-level ablations further show that the continuous target, rather than the exact scalar-shaping formula, was the principal additional signal. Separate fixed-policy tests obtained 0% success rates on two unseen maze layouts. These results indicate that continuous subgoal targets can encode task-specific route information in the source maze, while cross-layout transfer remains unresolved.

Li-Dong Sun, Ye Wang, Zhennan Fan et al. · 0 citations
Conference Aug 2026

An Obstacle Avoidance Path Planning Method Based on Improved OPSN for Robotic Arm

This paper proposes a three-dimensional obstacle avoidance path planning method for a single-arm manipulator based on an improved Optimization Problem Solving Network (OPSN). To address the difficulties caused by non-convex search spaces, complex obstacle constraints, and the poor performance of conventional swarm intelligence algorithms in narrow feasible regions, the end-effector trajectory is modeled as a polyline with fixed start and goal points and several intermediate waypoints. Path length, trajectory smoothness, and task-related height preference are jointly incorporated into the objective function, while workspace boundary constraints, obstacle safety distance constraints, and minimum height constraints are explicitly embedded into the network structure. In addition, an elite-initialization strategy is introduced to improve the original OPSN, whose initial inputs are purely random and cannot exploit useful historical information across restarts. The proposed strategy maintains exploration in the early stage and generates new initializations from an elite pool in the later stage through adaptive perturbation and weighted combination. Comparative experiments in three representative scenarios show that the improved OPSN achieves superior or competitive overall performance, especially in narrow-passage environments, where it exhibits stronger feasible-solution search capability and shorter planned paths.

Jianhan Fan, C. Peng, Jian-Xiao Zou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.