A fully built, integration-tested application that lets a user describe an investment goal in plain language and produces a live, broker-integrated portfolio recommendation from athree-phase reinforcement learning system, which generalize to other applied RL systems built on external, live data sources.
Abstract
Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- existing robo-advisors use static, rule-based allocation, and institutional-grade systems require account minimums and technology stacks unavailable to individual investors. We present a fully built, integration-tested application that closes this gap: a FastAPI backend and web dashboard that let a user describe an investment goal in plain language (e.g."I want steady growth but need to sell some shares next month for a down payment"), routes that goal to one of six investment mandates, and produces a live, broker-integrated portfolio recommendation from athree-phase reinforcement learning system -- a self-supervised cross-asset encoder, a Mixture-of-Experts (MoE) allocation policy with a learned intent router, and a lightweight LoRA adapter that personalizes recommendations from an individual's revealed brokerage behavior without retraining the shared model. The system is functionally complete and integration-tested end-to-end against a live brokerage API (Alpaca, paper-trading mode), including multi-user authentication, a trust first preview-before-apply confirmation flow, daily email digests, and an auditable action-integrity chain, but has not yet been opened to real end-users; we report this honestly as an emerging, pre-deployment application with a concrete path to full deployment, alongside 14-day walk-forward backtests (bootstrapped confidence intervals included) as preliminary, pre-deployment validation rather than production performance. We also report several practical engineering lessons -- silently-inactive integration paths, hanging third-party API calls, and the value of end-to-end empirical verification over trusting checkpoint metadata -- that we believe generalize to other applied RL systems built on external, live data sources.
This study introduces a Transformer-driven modeling approach that jointly performs product recommendation and personalized price-level assignment by modeling observed discount-level acceptance as an operational proxy for consumers’ willingness to pay (WTP).
Ali Mahdavian, Hadi Moradi, B. Bahrak· Journal of Theoretical and A...· 0 citations
The end-to-end treatment policy delivered a statistically significant $+7.20\% lift in the primary long-term-value metric, demonstrating the feasibility of production-scale causal optimization under business constraints.
Changshuai Wei, John Bencina, Phuc Nguyen et al.· 0 citations
This work offers a highly adaptable framework that successfully aligns multi-objective algorithmic trading with diverse, real-world human sustainability preferences and integrates a Preference Elicitation framework using Gaussian Processes.
Giovanni Dispoto, Marcello Restelli, Carmine Ventre· 0 citations
This work proposes a retrieval-augmented expert-switching framework that dynamically selects portfolio management experts based on their historical performance under similar market situations based on their historical performance under similar market situations.
Tax-loss harvesting demonstrates consistent benefits to long-term portfolio growth; yet implementing it efficiently often involves complex considerations that are specific to the holdings within that portfolio and the individual who owns it. We introduce a custom capital gains calculation engine and a RAG-retrieved vec...
Aryan Brar, Justin Du, Avery Lor et al.· 0 citations
Portfolio reinforcement learning (RL) commonly represents each action as a complete asset-weight vector, causing the action dimension and exploration difficulty to grow with the investment universe. This study proposes FrontierStep-RL, which replaces the direct N-dimensional action with two bounded variables: a frontie...