Who this is for. Researchers considering RL for an execution, market-making or allocation problem, and readers of the RL-trading literature. Assumes basic RL vocabulary (policy, reward, environment).

How to use it. Read the steps in order; each one assumes the previous. The “why here” line says what the step adds. Where a step is a tool, run it on your own numbers before moving on. The exemplar papers at the end are the archive’s highest-rigor papers for the path’s methods; read two of them with the checklists in hand.

Steps

  1. A Visual Map of Reinforcement Learning for Trading

    A Visual Map of Reinforcement Learning for Trading: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored.

    Why here: The lay of the land: which problems RL papers attack and how the archive scores them.

  2. How to Evaluate a Reinforcement-Learning Trading Paper

    A checklist for evaluating RL trading papers: environment leakage, reward hacking, the sample-efficiency problem in non-stationary markets, and where RL claims are actually credible.

    Why here: The checklist: seeds, environment realism, baselines, costs.

  3. Market Making: Inventory Risk, Adverse Selection, and Spread Capture

    The three forces of market making — spread capture, inventory risk, and adverse selection — with the Avellaneda-Stoikov intuition, the profitability identity, and the production failure catalog.

    Why here: The problem RL most plausibly helps with, and the static baseline it must beat.

  4. A Visual Map of Market-Making Research

    A Visual Map of Market-Making Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored.

    Why here: Market-making research mapped by score.

  5. Transaction Costs, Slippage, and Market Impact: From Paper Alpha to Tradeable Alpha

    How to model transaction costs in backtests: the four cost components, the square-root impact law, honest cost ranges by asset class, and the turnover arithmetic that kills most published strategies.

    Why here: Reward functions that ignore impact train agents that trade too much.

  6. Regime Dependence: When a Historical Edge Stops Working

    Regime dependence in trading strategies: why edges are conditional on market states, how to detect regime-carried backtests, decay vs regime-shift diagnosis, and what to do when an edge goes quiet.

    Why here: A policy trained on one regime is a bet on that regime.

  7. A Minimal ML Experiment-Tracking Stack for Quant Research

    Experiment tracking for quant ML: what to record, the finance-specific requirements (trial counts, temporal splits, leakage audits), tool tiers from SQLite to MLflow/W&B, and the minimal stack that suffices.

    Why here: Many seeds and many configs need bookkeeping.

  8. Quant Researcher Compute-Cost Calculator tool

    Interactive calculator: local GPU vs cloud GPU vs API costs for quant research. Editable assumptions, break-even hours, and a monthly budget for your whole stack.

    Why here: What the honest version of the experiment costs to run.

All paths: learning paths. Methods, with their own guides and exemplars: methods pages. Programmatic access to the papers: agents page.