Felipe Moret, Fabrizio Lillo · 2026-09-10
A plain-English AI summary of what this paper means for investors — generated on demand from the abstract.
Classical market-making strategies based on stochastic control, such as the Avellaneda-Stoikov and the Guéant-Lehalle-Fernandez-Tapia (GLFT) extension, provide closed-form quoting rules, but rest on assumptions that break down at realistic microstructure timescales. One of them is that order flow is stationary, while empirical evidence points to the existence of regimes, possibly associated with algorithmic execution of metaorders. In this case, existing methods provide negative PnL. In this paper, we develop a deep reinforcement-learning market maker (RLMM) - a Rainbow-style distributional DQN (C51) which is calibrated and tested in a zero-intelligence limit order book. We find that, in the stationary setting, RLMM outperforms GLFT across the entire observed risk-return frontier. The RLMM is more robust to flow asymmetry than GLFT, but, like any stationarily trained strategy, it still suffers large drawdowns from inventory saturation under persistent directional imbalance. Augmenting the state of RLMM with two auxiliary signals - a Bayesian online change-point filter over the directional flow bias and a queue-adjusted quote-exposure imbalance -restores profitability. A final scenario-bandit step that reweights low-return regime scenarios further improves performance under random-persistence and correlated-direction stress.
Go deeper: a full research-committee breakdown of this paper, its assumptions and failure modes, and how its method would apply to a specific ticker or your watchlist. See StockTools AI →
AI summary generated from the paper’s public abstract via arXiv; it may miss nuance — read the source before relying on it. Thank you to arXiv for its open-access interoperability; StockTools is not affiliated with arXiv, and all rights remain with the authors. Educational only, not financial advice.