Skip to content

Quantum-Inspired Portfolio Optimization Using Reinforcement Learning for Dynamic Stock Allocation

Aug 2026 · International Journal of Computational Science and Engineering Research · Vol 3, pp. 50 · 0 citations

TL;DR

A new Quantum-Inspired Portfolio Optimization (QIPO-RL) model is presented that combines quantum-inspired search techniques, an adaptive RL agent, and asset weights to create a framework for a Reinforcement Learning (RL) based dynamic stock allocation algorithm.

Abstract

Despite the dynamic nature of the market, large dimensionalities of asset space and varieties of financial products, portfolio optimization is a highly challenging problem. These traditional methods like Markowitz Mean-Variance Optimization and Risk Parity are highly static, and have not been able to adjust to the speed-of-change in market regimes. In this paper, the authors present a new Quantum-Inspired Portfolio Optimization (QIPO-RL) model that combines these interesting approaches to create a framework for a Reinforcement Learning (RL) based dynamic stock allocation algorithm. The framework blends quantum-inspired search techniques, an adaptive RL agent, and asset weights to maximize risk-adjusted asset returns, while maintaining assets in optimal allocation, and provides a way to rebalance a portfolio continuously in response to the changing market environment. The results from experiments were compared with state of the art baselines, which showed that QIPO-RL’s annual return is 16.8%, its Sharpe ratio is 1.61, and the maximum drawdown is 11.5% which is the best among all the competing methods. These findings support the synergism of using a global search method inspired by quantum computers in conjunction with an RL-based adaptation of decisions.

View source

Similar papers

Jul 2026

SciPhy Reinforcement Learning for Portfolio Optimization

The results demonstrate that the proposed framework successfully translates known signal quality into a robust, multi-period, and cost-aware allocation mechanism with strictly controlled volatility and turnover.

I. Halperin, A. Itkin · 0 citations
Open access 2026

A Novel Multi-Objective Quantum-Inspired Algorithm for Portfolio Optimization with Short-Selling in Real-World Market

: Portfolio optimization is inherently a multi-objective problem that aims to maximize expected return while minimizing investment risk, while also facing exponential growth in the search space and increasing market complexity. Existing multi-objective optimization approaches often struggle to balance convergence and diversity, particularly under realistic trading conditions such as short-selling. To address these challenges, this paper proposes a novel Multi-objective Quantum-inspired Tabu Search (MoQTS) framework for portfolio optimization with short-selling strategies. The proposed method incorporates a quantum-inspired superposition mechanism to enhance global exploration and introduces an entanglement-driven neighborhood search strategy that systematically generates structured local perturbations by modifying one or two asset-selection states. This mechanism enables effective exploration of the neighborhood of non-dominated solutions, thereby improving both convergence accuracy and solution diversity. In addition, a trend ratio (TR)-based evaluation model is adopted to jointly capture return and risk dynamics under real-world market fluctuations. Experiments are conducted on the U.S. stock market using Dow Jones Industrial Average (DJIA) data from 2013 to 2025. The proposed MoQTS is compared with several state-of-the-art multi-objective algorithms, including NSGA-II, MOEA/D, SMS-EMOA, and MOPSO. Experimental results demonstrate that MoQTS can obtain high-quality Pareto-optimal solutions with strong convergence and diversity performance. Results over multiple independent periods further support the effectiveness and robustness of MoQTS. In addition, computational cost analysis shows that MoQTS substantially reduces the number of evaluations and execution time compared with the comparison algorithms while maintaining high-quality Pareto-optimal solutions.

Yun-Ting Lai, Ming Chang, Yao-Hsin Chou · 0 citations
Conference Jul 2026

Practical Quantum Portfolio Optimization Under Discrete Lot Constraints with CVaR-Based Evaluation for the Taiwan Stock Market

This paper proposes a CVaR-based quantum portfolio optimization framework designed to address discrete market constraints. Unlike traditional models that often assume continuous asset allocation or normal return distributions, the proposed approach utilizes Conditional Value-at-Risk (CVaR) as the objective function within the hybrid optimization loop of a gate-based Quantum Approximate Optimization Algorithm (QAOA) to better manage extreme tail risk in realistic financial portfolios. We formulate the portfolio selection and discrete constraints as a Knapsack-style Quadratic Unconstrained Binary Optimization (QUBO) model, explicitly incorporating the “one-lot” (1,000 shares) trading convention common in the Taiwan Stock Exchange (TWSE). Experimental results on small-scale TWSE instances indicate that the proposed framework achieves lower CVaR values than standard expectation-based QAOA, albeit with a moderate reduction in expected return. These findings provide preliminary evidence that quantum optimization can support risk-aware portfolio selection under discrete trading constraints.

Yun-Ching Lu, Tzung-Her Chen · 0 citations
Open access Jul 2026

Ant colony intelligence optimization of asset portfolios under constraints—dynamic risk adjustment and Sharpe ratio maximization

This study addresses the challenge of constrained portfolio optimization by proposing a novel framework based on an enhanced Ant Colony Optimization (ACO) algorithm. Building upon the Markowitz mean-variance foundation, we propose a hybrid framework that integrates Ant Colony Optimization (ACO) for discrete asset selection under cardinality constraints with quadratic programming for optimal weight allocation subject to no-short-selling and budget constraints. Dynamic risk estimation is achieved by employing EWMA and GARCH(1,1) models on rolling windows to update the covariance matrix during optimization. The latter is achieved by employing EWMA and GARCH(1,1) models on rolling windows to update the covariance matrix during optimization. We introduce key enhancements to the ACO algorithm, such as an elite retention strategy and an adaptive pheromone update mechanism linked to asset weights, to improve convergence and solution quality under complex constraints. Empirical validation is conducted using historical daily returns from a stratified sample of 20 S&P 500 stocks. Results demonstrate that the enhanced ACO significantly outperforms benchmark algorithms (PSO and GA), yielding portfolios with higher Sharpe ratios (average improvement of 10–15% on test set) and superior risk control (reduced maximum drawdown by 5–7%). Robust performance is maintained across both bull and bear market regimes. This research provides an effective heuristic optimization approach for quantitative multi-factor strategy design, offering practical value for real-world asset allocation under constraints.

Yu-Qun Cao, Lu Yu, Qunqun Cao · 0 citations
Open access Sep 2026

FrontierStep-RL: Fixed-Dimensional Structured Actions for Transaction-Cost-Aware Portfolio Reinforcement Learning

Portfolio reinforcement learning (RL) commonly represents each action as a complete asset-weight vector, causing the action dimension and exploration difficulty to grow with the investment universe. This study proposes FrontierStep-RL, which replaces the direct N-dimensional action with two bounded variables: a frontier coordinate and a rebalancing step. At each decision date, rolling estimates of expected returns and covariance define a regularized efficient frontier. A cost–risk-aware coordinate organizes the frontier using normalized local changes in predicted volatility and one-way turnover. The coordinate selects a frontier-supported target portfolio, while the step controls how far the pre-trade portfolio moves toward that target. We evaluate FrontierStep-RL on FF49, FF100, and FNSPID-50 against traditional strategies, controlled direct-weight RL policies, and recent portfolio-management methods. FrontierStep-RL achieves net Sharpe ratios of 0.75 on both FF49 and FNSPID-50 while maintaining comparatively low volatility, drawdown, and turnover. In the 100-asset setting, it achieves 0.68, compared with 0.52 for the strongest direct-weight baseline, and completes all runs. At a transaction cost of 50 basis points, it retains net Sharpe ratios of 0.559 and 0.568. The results support fixed-dimensional target selection and controlled execution for scalable, transaction-cost-aware portfolio RL.

Hou-Yu Zou, Hui Li, Feng Xue et al. · 0 citations
Open access Aug 2026

Benchmarking deep reinforcement learning and classical models for portfolio optimization across market efficiency regimes

A key puzzle in finance is why algorithmic traders with advanced neural models sometimes fail to beat simple traditional strategies, while in other cases they clearly outperform them. This study argues that such variation depends on how information is reflected in market prices. When markets are highly efficient, price dynamics are stable and structured. This environment is well suited for deep reinforcement learning, which can learn adaptive allocation patterns. In less efficient markets, price dynamics are noisier and more unstable, which makes it harder for data intensive models to perform consistently. To examine this idea, we analyse forty-five Nifty 50 stocks using a time varying Fuzzy Market Inefficiency Measure and group them into three efficiency levels. Within each cluster, four deep reinforcement learning models are compared with six traditional portfolio strategies under the same conditions. The results show that deep learning models perform best in highly efficient markets where signals are weak but consistent. In moderately and least efficient markets, traditional strategies often achieve similar or better returns. However, deep learning models still provide better control over downside risk in less efficient environments. Overall, the findings offer valuable insights for portfolio managers and investors by supporting efficiency-based portfolio allocation, enhanced risk management, and adaptive investment strategies across varying market conditions.

H. Sahu, Avishek Bhandari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.