Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi). Constant product markets with concentrated liquidity, such as UniswapV3, are now a well-established design. In these markets, liquidity providers (LPs) face a sequential decision problem: they must decide when to rebalance their positions and which price ranges to allocate capital to as market conditions evolve. We formulate dynamic liquidity provision as a stochastic impulse control problem and use reinforcement learning (RL) to solve it, focusing on providing interpretable solutions. We show that learned policies exhibit rich state-dependent behaviour, allocating liquidity according to mispricing, rebalancing costs, uncertainty, inventory exposure, and heterogeneous risk preferences. These behaviours help compress the left tail of the Profit and Loss (PnL) distribution and avoid catastrophic outcomes under high uncertainty. Finally, we benchmark the RL agents against baseline and sophisticated agents from the AMM microstructure literature and analyse their performance.
Modern electronic markets are increasingly populated by adaptive artificial intelligence (AI)–driven trading systems that continuously learn from market outcomes and from one another. This article studies a specific form of algorithmic interaction risk. We simulate multiple market-making agents using tabular Q-learning to optimize quote aggressiveness in the presence of noise traders. Although the agents do not communicate or optimize a joint objective, repeated interaction leads to persistent spread widening, increased dealer profitability, and deterioration in transaction-cost proxies faced by liquidity demanders. Using memory-1 Q-learning agents that compete only on spread, we find that two-dealer markets converge to outcomes that close 78% of the gap between the static Bertrand–Nash equilibrium and joint monopoly, with 97.5% of the joint-action mass on the diagonal (both liquidity providers quoting the same spread). This article develops several observable diagnostics for detecting such behavior, including spread persistence, quote clustering, and forced-deviation response tests. Unlike prior work focused on informed trading and price discovery in Kyle-style environments, our framework emphasizes quote competition, liquidity provision, and execution quality in electronic markets. The results suggest that adaptive interaction among AI trading systems may create new forms of algorithmic vulnerability relevant for exchanges, trading venues, execution desks, and market-surveillance teams.
Sudip Gupta· The Journal of Financial Dat...· 0 citations
When economic structures and market dynamics shift, classic portfolio rebalancing algorithms often suffer from unstable and degraded performance. To improve the return and robustness of portfolio management, we explore reinforcement learning (RL) and propose Scenario-Context Rollout (SCR), a macroeconomics-guided feedback mechanism to produce a distribution of next-day joint returns under potential economic shocks. However, doing so faces new challenges, as history will never tell what would have happened differently. As a result, incorporating scenario-based rewards from rollouts introduces a reward-transition mismatch in temporal-difference (TD) learning, destabilizing RL critic training. We theoretically analyze this problem and show that combining scenario-scored rewards with tape-realized transitions induces a hybrid fixed point. Guided by this analysis, we construct a counterfactual next state using the SCR continuations and augment the critic agent's bootstrap target. Doing so stabilizes the learning and provides a viable bias-variance tradeoff. In out-of-sample evaluations across 31 distinct universes of U.S. equity and ETF portfolios, our method improves Sharpe ratio by up to 76% and reduces maximum drawdown by up to 53% compared with classic and RL-based portfolio rebalancing baselines.
Vanya Priscillia Bendatu, Yao Lu· Proceedings of the 32nd ACM...· 0 citations
This paper examines the forecasting of liquidity dynamics in European stock markets by means of traditional econometric models and machine learning techniques. It uses daily data for the DAX, CAC 40, FTSE 100, FTSE MIB, and IBEX 35 over 2010–2026, liquidity being measured by the logarithmic Amihud illiquidity indicator. The empirical framework compares ARIMA models and a dynamic panel specification with Random Forest, Extreme Gradient Boosting (XGBoost), and Support Vector Regression (SVR) within a common rolling one-step-ahead forecasting framework. The results show that liquidity is highly persistent and that the dynamic panel model achieves the lowest forecast errors, although Diebold–Mariano tests indicate no significant predictive advantage over the leading machine learning models. SHAP analysis reveals that trading activity, lagged liquidity, and market uncertainty are the main determinants of liquidity forecasts. The findings highlight the complementary role of explainable machine learning in empirical finance.
Veni Arakelia, G. Caporale, Mirto M Gasparinatou et al.· CESifo working papers· 0 citations
Hedging a derivative position under transaction costs and market frictions requires a trading rule that adapts to changing conditions. Deep hedging trains a neural policy for this task but policy training does not determine whether a trading desk can afford to run the policy. We apply robust hedging valuation adjustment (HVA) as a post-training valuation-adjustment layer that evaluates tracking-loss CVaR together with explicit funding and margin add-ons. The funding and margin add-ons share the same KL uncertainty set as HVA. For each policy, a single common-stress tilt computes HVA, funding and margin jointly and a trading desk can get one internally consistent reserve instead of the three separately. We compare classical hedge policies with learned hedge specifications across three market environments with different liquidity. No single specification dominates in every market. Under the strict tracking-risk budget, gamma-wide classical bands are selected in High and Middle Liquidity while sparse learned execution is selected in Low Liquidity. At looser validation budgets wider classical bands are generally selected.
A closed-loop multi-agent decision framework that introduces prompt-level learning as a scalable alternative to full model retraining and highlights the potential of prompt-level adaptation for building robust and autonomous financial decision systems.
Kandarp Mukeshkumar Sharda, Aliyu Sani Sambo· NLP & Big Data· 0 citations
Classical put-overlays have long been treated as a reliable hedge against tail risk but the market conditions of 2025 expose their limits in ways that theory didn’t fully anticipate. This paper examines where these strategies break down: in markets defined by elevated volatility, wide bid-ask spreads, and structural frictions that quietly erode the protection investors thought they’d bought. Traditional hedging methods, which rely on the systematic purchase of out-of-the-money (OTM) options, are increasingly hampered by high premiums and reliance on static volatility assumptions. During the significant geopolitical disruptions of 2025, most notably the “Tariff Shock” of April, these traditional models proved inadequate, suffering from “premium bleed” and an inability to account for discontinuous market gaps. To address these systemic vulnerabilities, we formulate the hedging process as a stochastic control problem and implement a Deep Hedging framework utilising Deep Reinforcement Learning (DRL). Optimised specifically for Expected Shortfall (ES), the model utilises Long Short-Term Memory (LSTM) layers to process multi-dimensional state vectors, including implied volatility skew and realised turbulence. Our empirical results demonstrate that this AI-optimised policy achieves a 35% reduction in hedging costs while simultaneously improving tail risk protection and draw-down resilience. Notably, the DRL agent exhibits anticipatory behaviour, transitioning from a reactive to a predictive paradigm by adjusting hedge positions prior to observable volatility spikes. These findings suggest that in structurally incomplete markets, optimal risk management has evolved from simple insurance into a process of continuous, regime-aware policy optimisation. This study provides a robust framework for institutional solvency in an era defined by non-linear correlations and rapid liquidity decay.
Y. Chakrabarti· American Journal of Financia...· 0 citations