Paper: arXiv 2610.03598

Authors: Siu Tung Wong, Carlo Campajola

Abstract

Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker–trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker’s continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO–FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO–FFNN and PPO–LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker’s observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes (2.22%) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

Complexity vs Empirical Score

  • Math Complexity: 7.5/10
  • Empirical Rigor: 7.0/10
  • Quadrant: Holy Grail — high math complexity, high empirical rigor

Why this score: This paper demonstrates a strong blend of advanced mathematical modeling for the broker-trader game and rigorous empirical testing of RL agents against analytical benchmarks. The novelty lies in using analytically solved models to diagnose and guide PPO, offering a unique perspective on RL’s application in finance.

Research Flowchart

  flowchart TD
    A[Research Goal: Can RL exploit solved financial models?] --> B{Methodology: PPO in Broker-Trader Game};
    B --> C[Data/Inputs: Analytically Solved Continuous-Time Broker-Trader Game, Stochastic Uninformed Order Flow];
    C --> D{Computational Processes: PPO-FFNN, PPO-LSTM, Monte Carlo Diagnostics, Supervised Learning, Causal Certainty-Equivalent Controller};
    D --> E[Outcomes: PPO Inaccuracy with Stochastic Flow; Critic Fails to Rank Actions; CE Controller Outperforms PPO; PPO Adapts to Cost Changes];