Who this is for. Practitioners and students who can train a model and want to judge published ML results, or avoid publishing a bad one. Assumes familiarity with supervised learning.
How to use it. Read the steps in order; each one assumes the previous. The “why here” line says what the step adds. Where a step is a tool, run it on your own numbers before moving on. The exemplar papers at the end are the archive’s highest-rigor papers for the path’s methods; read two of them with the checklists in hand.
Steps
- How to Evaluate a Machine-Learning Trading Paper
A checklist for evaluating ML trading papers: baseline honesty, leakage-prone pipelines, accuracy-vs-P&L confusion, seed variance, and the questions that separate signal from citation bait.
Why here: The master checklist; everything below expands one of its items.
- Look-Ahead Bias and Point-in-Time Data: The Complete Taxonomy
Every way future information leaks into backtests: restated fundamentals, index membership, same-bar execution, timestamp semantics, and LLM training-data leakage — with detection tests for each.
Why here: Feature construction is where ML papers leak the future.
- Combinatorially Symmetric Cross-Validation (CSCV) Explained
CSCV and the Probability of Backtest Overfitting (PBO) explained: the block-combination construction, pseudocode, how to read PBO, and the method's honest limitations.
Why here: Cross-validation that respects time and counts trials.
- P-Hacking in Financial Research: The Practices, the Tells, the Fixes
How p-hacking works in finance: the seven questionable research practices, the tells visible in published papers, and the reader-side corrections that keep you from funding other people's noise.
Why here: The tells of a result selected from many.
- Statistical Significance vs Economic Significance in Trading Research
Why t-statistics mislead in finance: multiple testing and the factor zoo, the t > 3 hurdle, economic magnitude after costs, and the two-by-two matrix for judging any empirical result.
Why here: Accuracy above 50% is not a strategy.
- Transaction Costs, Slippage, and Market Impact: From Paper Alpha to Tradeable Alpha
How to model transaction costs in backtests: the four cost components, the square-root impact law, honest cost ranges by asset class, and the turnover arithmetic that kills most published strategies.
Why here: High-turnover ML signals live or die here.
- A Minimal ML Experiment-Tracking Stack for Quant Research
Experiment tracking for quant ML: what to record, the finance-specific requirements (trial counts, temporal splits, leakage audits), tool tiers from SQLite to MLflow/W&B, and the minimal stack that suffices.
Why here: Record the trial count so the haircut is honest.
- The Deflated Sharpe Ratio, Explained with Real Numbers
How the deflated Sharpe ratio corrects for multiple testing, fat tails, and short samples: the intuition, the formulas, and the worked example — 100 random backtests produce a Sharpe ≈ 2.5 by luck alone.
Why here: Apply the haircut.
- Deflated Sharpe Ratio & Minimum Track Record Calculator tool
Interactive deflated Sharpe ratio calculator: probabilistic Sharpe ratio, expected maximum Sharpe under N zero-skill trials, DSR, and minimum track record length — with skew, kurtosis, and sample length. Runs in your browser.
Why here: Run the numbers for the paper you are reading.
- A Visual Map of Machine Learning in Asset Pricing
A Visual Map of Machine Learning in Asset Pricing: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored.
Why here: Where the ML-in-asset-pricing literature stands.
Related
All paths: learning paths. Methods, with their own guides and exemplars: methods pages. Programmatic access to the papers: agents page.