Who this is for. Practitioners and students who can train a model and want to judge published ML results, or avoid publishing a bad one. Assumes familiarity with supervised learning.

How to use it. Read the steps in order; each one assumes the previous. The “why here” line says what the step adds. Where a step is a tool, run it on your own numbers before moving on. The exemplar papers at the end are the archive’s highest-rigor papers for the path’s methods; read two of them with the checklists in hand.

Steps

  1. How to Evaluate a Machine-Learning Trading Paper

    A checklist for evaluating ML trading papers: baseline honesty, leakage-prone pipelines, accuracy-vs-P&L confusion, seed variance, and the questions that separate signal from citation bait.

    Why here: The master checklist; everything below expands one of its items.

  2. Look-Ahead Bias and Point-in-Time Data: The Complete Taxonomy

    Every way future information leaks into backtests: restated fundamentals, index membership, same-bar execution, timestamp semantics, and LLM training-data leakage — with detection tests for each.

    Why here: Feature construction is where ML papers leak the future.

  3. Combinatorially Symmetric Cross-Validation (CSCV) Explained

    CSCV and the Probability of Backtest Overfitting (PBO) explained: the block-combination construction, pseudocode, how to read PBO, and the method's honest limitations.

    Why here: Cross-validation that respects time and counts trials.

  4. P-Hacking in Financial Research: The Practices, the Tells, the Fixes

    How p-hacking works in finance: the seven questionable research practices, the tells visible in published papers, and the reader-side corrections that keep you from funding other people's noise.

    Why here: The tells of a result selected from many.

  5. Statistical Significance vs Economic Significance in Trading Research

    Why t-statistics mislead in finance: multiple testing and the factor zoo, the t > 3 hurdle, economic magnitude after costs, and the two-by-two matrix for judging any empirical result.

    Why here: Accuracy above 50% is not a strategy.

  6. Transaction Costs, Slippage, and Market Impact: From Paper Alpha to Tradeable Alpha

    How to model transaction costs in backtests: the four cost components, the square-root impact law, honest cost ranges by asset class, and the turnover arithmetic that kills most published strategies.

    Why here: High-turnover ML signals live or die here.

  7. A Minimal ML Experiment-Tracking Stack for Quant Research

    Experiment tracking for quant ML: what to record, the finance-specific requirements (trial counts, temporal splits, leakage audits), tool tiers from SQLite to MLflow/W&B, and the minimal stack that suffices.

    Why here: Record the trial count so the haircut is honest.

  8. The Deflated Sharpe Ratio, Explained with Real Numbers

    How the deflated Sharpe ratio corrects for multiple testing, fat tails, and short samples: the intuition, the formulas, and the worked example — 100 random backtests produce a Sharpe ≈ 2.5 by luck alone.

    Why here: Apply the haircut.

  9. Deflated Sharpe Ratio & Minimum Track Record Calculator tool

    Interactive deflated Sharpe ratio calculator: probabilistic Sharpe ratio, expected maximum Sharpe under N zero-skill trials, DSR, and minimum track record length — with skew, kurtosis, and sample length. Runs in your browser.

    Why here: Run the numbers for the paper you are reading.

  10. A Visual Map of Machine Learning in Asset Pricing

    A Visual Map of Machine Learning in Asset Pricing: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored.

    Why here: Where the ML-in-asset-pricing literature stands.

All paths: learning paths. Methods, with their own guides and exemplars: methods pages. Programmatic access to the papers: agents page.