Skip to content

A comparison between Thompson sampling and greedy algorithm in portfolio selection

Abstract

Portfolio optimization is a decision-making problem that allocates assets to achieve an optimal return–risk trade-off. The classical framework assumes that return and risk parameters are known; in practice, however, these parameters must be estimated from data, introducing estimation risk beyond the original model. Integrating parameter learning with portfolio choice naturally leads to a sequential decision-making framework based on reinforcement learning. This study examines sequential portfolio selection under parameter uncertainty within a Bayesian portfolio framework. It evaluates whether exploration through Thompson sampling improves performance relative to an exploration-free greedy strategy when asset returns are observable and portfolio weights do not influence the return-generating process. The analysis includes a simulation study and an empirical application using daily returns of eight stocks from the SET100 index. Portfolio weights are updated sequentially as posterior beliefs about expected returns evolve over time. Performance is assessed using cumulative return, Sharpe ratio, and regret. Simulation results show that Thompson sampling does not outperform the greedy algorithm across performance metrics; both approaches generate nearly identical outcomes. Empirical results using real data indicate that Thompson sampling occasionally slightly outperforms the greedy algorithm, but the differences are not economically significant. Overall, the findings suggest that in standard passive investment settings, the additional complexity of exploration-based reinforcement learning is not justified. Simpler estimation-based approaches deliver comparable performance.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.