Skip to content
Open access

FrontierStep-RL: Fixed-Dimensional Structured Actions for Transaction-Cost-Aware Portfolio Reinforcement Learning

Hou-Yu Zou Hui Li Feng Xue Tian-Hao Yuan
Sep 2026 · Mathematics · 0 citations · 51 references

Abstract

Portfolio reinforcement learning (RL) commonly represents each action as a complete asset-weight vector, causing the action dimension and exploration difficulty to grow with the investment universe. This study proposes FrontierStep-RL, which replaces the direct N-dimensional action with two bounded variables: a frontier coordinate and a rebalancing step. At each decision date, rolling estimates of expected returns and covariance define a regularized efficient frontier. A cost–risk-aware coordinate organizes the frontier using normalized local changes in predicted volatility and one-way turnover. The coordinate selects a frontier-supported target portfolio, while the step controls how far the pre-trade portfolio moves toward that target. We evaluate FrontierStep-RL on FF49, FF100, and FNSPID-50 against traditional strategies, controlled direct-weight RL policies, and recent portfolio-management methods. FrontierStep-RL achieves net Sharpe ratios of 0.75 on both FF49 and FNSPID-50 while maintaining comparatively low volatility, drawdown, and turnover. In the 100-asset setting, it achieves 0.68, compared with 0.52 for the strongest direct-weight baseline, and completes all runs. At a transaction cost of 50 basis points, it retains net Sharpe ratios of 0.559 and 0.568. The results support fixed-dimensional target selection and controlled execution for scalable, transaction-cost-aware portfolio RL.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.