Reinforcement learning promises to learn the policy directly, skipping the forecast-then-optimize split, and it is a natural fit for sequential problems like optimal execution and market making where the environment can be simulated. Most of the archive’s RL papers instead train an agent on a single historical price series and report its backtest, which is the setting where RL is least defensible: one non-stationary trajectory, a reward that is a noisy function of a few hundred decisions, and a model with millions of parameters.

What to check when reading. Seeds and variance first: a credible paper trains many seeds and reports the distribution, not the best run. Then the environment: does the simulator model costs, latency, and the agent’s own impact, and was it fitted on the test period? Then the baseline: a tuned static rule (TWAP for execution, a Stoikov quote for market making, equal-weight for allocation) often matches the RL result. The guide on evaluating RL papers is a checklist for exactly this.