Who this is for. Researchers considering RL for an execution, market-making or allocation problem, and readers of the RL-trading literature. Assumes basic RL vocabulary (policy, reward, environment).
How to use it. Read the steps in order; each one assumes the previous. The “why here” line says what the step adds. Where a step is a tool, run it on your own numbers before moving on. The exemplar papers at the end are the archive’s highest-rigor papers for the path’s methods; read two of them with the checklists in hand.
Steps
- A Visual Map of Reinforcement Learning for Trading
A Visual Map of Reinforcement Learning for Trading: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored.
Why here: The lay of the land: which problems RL papers attack and how the archive scores them.
- How to Evaluate a Reinforcement-Learning Trading Paper
A checklist for evaluating RL trading papers: environment leakage, reward hacking, the sample-efficiency problem in non-stationary markets, and where RL claims are actually credible.
Why here: The checklist: seeds, environment realism, baselines, costs.
- Market Making: Inventory Risk, Adverse Selection, and Spread Capture
The three forces of market making — spread capture, inventory risk, and adverse selection — with the Avellaneda-Stoikov intuition, the profitability identity, and the production failure catalog.
Why here: The problem RL most plausibly helps with, and the static baseline it must beat.
- A Visual Map of Market-Making Research
A Visual Map of Market-Making Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored.
Why here: Market-making research mapped by score.
- Transaction Costs, Slippage, and Market Impact: From Paper Alpha to Tradeable Alpha
How to model transaction costs in backtests: the four cost components, the square-root impact law, honest cost ranges by asset class, and the turnover arithmetic that kills most published strategies.
Why here: Reward functions that ignore impact train agents that trade too much.
- Regime Dependence: When a Historical Edge Stops Working
Regime dependence in trading strategies: why edges are conditional on market states, how to detect regime-carried backtests, decay vs regime-shift diagnosis, and what to do when an edge goes quiet.
Why here: A policy trained on one regime is a bet on that regime.
- A Minimal ML Experiment-Tracking Stack for Quant Research
Experiment tracking for quant ML: what to record, the finance-specific requirements (trial counts, temporal splits, leakage audits), tool tiers from SQLite to MLflow/W&B, and the minimal stack that suffices.
Why here: Many seeds and many configs need bookkeeping.
- Quant Researcher Compute-Cost Calculator tool
Interactive calculator: local GPU vs cloud GPU vs API costs for quant research. Editable assumptions, break-even hours, and a monthly budget for your whole stack.
Why here: What the honest version of the experiment costs to run.
Related
All paths: learning paths. Methods, with their own guides and exemplars: methods pages. Programmatic access to the papers: agents page.