# thequant.space — full text of guides and methodology > Companion to https://thequant.space/llms.txt . This file holds the complete text of the scoring methodology, the agent guide, every practitioner guide, every tool page and every topic-hub introduction. Dynamic paper lists are omitted here; the scored corpus is at https://thequant.space/data/papers.json . Free for personal and research use with attribution to thequant.space. ---------------------------------------------------------------------- # Scoring Methodology URL: https://thequant.space/score-guide/ Description: A framework for quantitative research classification and strategic navigation. # Scoring Methodology To navigate the high-volume landscape of quantitative finance, `thequant.space` utilizes a dual-axis scoring framework. This system decomposes complex ArXiv research into two primary dimensions: **Theoretical Complexity** and **Empirical Rigor**. --- ## The Classification Framework ### 1. Math Complexity This metric quantifies the "theoretical overhead" required to parse the paper. We analyze the density of mathematical notation and the sophistication of the underlying machinery. * **High Score (7–10):** Heavy reliance on **Stochastic Calculus**, **Partial Differential Equations (PDEs)**, bespoke optimization proofs, or advanced topology. * **Low Score (1–3):** Primarily descriptive statistics, high-level conceptual frameworks, or basic linear algebra. ### 2. Empirical Rigor This metric assesses the "path to implementation." We look for signals that the strategy has been stress-tested against the realities of the market. * **High Score (7–10):** Extensive backtesting on high-fidelity data (Tick-level, WRDS, Bloomberg), clear mention of transaction costs, and rigorous out-of-sample validation. * **Low Score (1–3):** Use of "toy" datasets, synthetic data, or purely logical derivations without historical verification. --- ## Strategic Quadrants By plotting these scores, we categorize research into four distinct strategic "vibes." This allows you to filter papers based on your current objective—whether it's deep R&D or immediate alpha generation. | Quadrant | Designation | Strategic Utility | | :--- | :--- | :--- | | **High Math + High Rigor** | **Holy Grail** | Institutional-grade research. Validated theory ready for sophisticated production environments. | | **High Math + Low Rigor** | **Lab Rats** | The "Research Frontier." Highly innovative math that lacks empirical testing—high potential for "hidden" alpha. | | **Low Math + High Rigor** | **Street Traders** | Applied Alpha. Robust, data-driven strategies that prioritize execution and simplicity over theoretical flair. | | **Low Math + Low Rigor** | **Philosophers** | Market Meta-Analysis. Conceptual frameworks and "think pieces" that shape broad market perspectives. | --- ## Current Library Distribution | Quadrant | Density | Primary Sources | | :--- | :---: | :--- | | **Holy Grail** | -- | J. of Finance, Quant. Finance | | **Lab Rats** | -- | ArXiv (Math.OC, Stat.ML) | | **Street Traders** | -- | Industry Whitepapers, SSRN | | **Philosophers** | -- | Commentary, Meta-Research | --- ## The Strategic Map > **Pro-Tip: Identifying "Alpha Clusters"** > Watch for clusters in the **Lab Rats** quadrant. These represent theoretical breakthroughs that have not yet been commoditized by the broader market. Bridging the gap from *Lab Rat* to *Street Trader* is where individual quants find their edge. --- *Disclaimer: Scoring is performed via automated semantic analysis. If you believe a paper has been misclassified, please open an issue.* ---------------------------------------------------------------------- # For AI Agents and Programmatic Users URL: https://thequant.space/agents/ Description: How to use thequant.space from an AI agent or a script: machine-readable data files, field semantics, llms.txt, citation format, and usage terms for the scored quant-finance research corpus. thequant.space is built to be read by software as well as by people. If you are an AI assistant, a research agent, or a script, this page is the contract: where the data is, what the fields mean, and what you may do with it. Nothing here needs a key, a login, or a browser. ## Entry points | What | URL | Notes | |---|---|---| | Site map for language models | [/llms.txt](/llms.txt) | What the site contains, how scoring works, every topic, guide and tool with a one-line description. Start here. | | Full text of the evergreen pages | [/llms-full.txt](/llms-full.txt) | The scoring methodology, every guide, every tool page and every topic introduction in one plain-text file. | | All scored papers | [/data/papers.json](/data/papers.json) | One record per paper, every field (about 4 MB; served gzip-compressed). Rebuilt daily. | | One topic hub | `/data/topics/.json` | Same record shape, 70 KB to 1 MB per file. Slugs and counts are listed inside `papers.json` under `topics`, and on [/topics/](/topics/). | | A single paper | `/flowcharts//` | HTML page with a `ScholarlyArticle` JSON-LD block in the head carrying the scores, plus a BibTeX entry at the bottom. The `page_url` field in the JSON points here. | | New pages | [/index.xml](/index.xml) | RSS, newest 50 pages, full text. | | Everything, for crawling | [/sitemap.xml](/sitemap.xml) | | The CSV version of the corpus is on the [dataset page](/dataset/) for people who prefer a spreadsheet. The JSON above is the same data. ## Record fields | Field | Meaning | |---|---| | `paper_id` | arXiv id (for example `2409.01234`) or `ssrn-` | | `title`, `paper_date` | Title and the paper's own date (arXiv submission or SSRN posting), `YYYY-MM-DD` | | `authors` | List of names; empty when the pipeline did not capture them (older SSRN items mostly) | | `topics` | List of [topic hub](/topics/) slugs assigned by the classifier; a paper can sit in several hubs | | `methods` | Method labels: `econometrics-time-series`, `stochastic-calculus`, `optimization`, `deep-learning`, `machine-learning`, `network-graph`, `nlp-llm`, `simulation`, `econophysics-complexity`, `reinforcement-learning`, `game-theory`, `bayesian`, `causal-inference`, `quantum` | | `asset_classes` | Asset-class labels from the classifier (equities, crypto, options, fixed income, commodities-energy, ...) | | `paper_type` | `empirical`, `theoretical`, `survey-review`, `methodology`, and similar classifier labels | | `primary_category` | Primary arXiv category, for example `q-fin.CP` | | `math_complexity` | 0 to 10. Theoretical overhead: 7+ means stochastic calculus, PDEs or bespoke proofs; 1 to 3 means descriptive statistics or conceptual frameworks | | `empirical_rigor` | 0 to 10. Path to implementation: 7+ means high-fidelity data, transaction costs and out-of-sample validation; 1 to 3 means toy or synthetic data, or no data | | `quadrant` | `Holy Grail` (high math, high rigor), `Street Traders` (low math, high rigor), `Lab Rats` (high math, low rigor), `Philosophers` (low math, low rigor). The split is at 5 on each axis | | `hub_score` | `0.6 × empirical_rigor + 0.4 × math_complexity`; the default ranking used across the site | | `code_url` | Public repository link when the paper gives one | | `doi`, `journal_ref` | When known; mostly null for preprints | | `paper_url` | The paper itself (arXiv PDF or SSRN abstract page) | | `page_url` | The paper's page on this site: summary, score rationale, research flowchart | | `tags` | Keywords extracted from the abstract | The scores are LLM-assisted judgements against a fixed rubric, not peer review. Read the [methodology](/score-guide/) before building conclusions on them, and treat a single paper's score as a prior, not a verdict. Papers are occasionally re-scored, so fetch fresh data rather than caching for months. ## Typical uses - **Research a topic.** Fetch the topic shard, filter `empirical_rigor >= 7`, sort by `paper_date` descending, read the `page_url` pages for the top hits. The page gives the abstract, the score rationale and the flowchart; the paper link gives the source. - **Learn a method.** Filter `papers.json` on `methods` for the method, keep `paper_type` in `survey-review` or `methodology` for the overview, then the highest `hub_score` empirical papers as worked examples. The [guides](/guides/) cover the evaluation side: backtest overfitting, walk-forward testing, the deflated Sharpe ratio, transaction costs, point-in-time data. - **Find reproducible work.** Filter `code_url` not null, or read [/papers-with-code/](/papers-with-code/). Pair with the [replication guide](/guides/reproduce-quant-research/) and the [replication checklist](/downloads/replication-checklist.md). - **Audit the scores.** The whole table is public; `math_complexity` and `empirical_rigor` by `topics` or by year is a one-line groupby. ## How to cite Cite the paper itself with the BibTeX entry on its page. When you use a score, a quadrant or a flowchart, attribute it as: > thequant.space, "", math complexity , empirical rigor , . . Scores are LLM-assisted rubric judgements; methodology at https://thequant.space/score-guide/. For the dataset as a whole: "thequant.space scored quant-finance research corpus, https://thequant.space/data/papers.json, retrieved ." ## Terms - The data files, scores, summaries and flowcharts are free for personal and research use with attribution to thequant.space, as above. - The papers themselves belong to their authors and publishers. We link to them; we do not redistribute them. - Crawling is welcome. The whole site is static; fetch `papers.json` once rather than 5,000 paper pages when you only need the table. There is no rate limit beyond Cloudflare's defaults. - Commercial redistribution of the dataset, or embedding it in a paid product, needs a conversation first: see [about](/about/) for contact. ## What is not here yet There is no query API and no MCP server yet; filtering is done client-side on the JSON. Both are planned. If you build something on this data, say so on the [about](/about/) page's contact route and we will link it. ---------------------------------------------------------------------- # About URL: https://thequant.space/about/ Description: What thequant.space is, how the paper pipeline works, and how to get in touch. # About thequant.space **Using the site from an AI agent or a script?** The machine-readable entry points (llms.txt, the full corpus as JSON, field definitions, citation format, usage terms) are on the [agents page](/agents/). thequant.space turns the firehose of quantitative finance research into something you can actually navigate. We process papers from **arXiv's q-fin sections and SSRN**, and for each one we publish: - A **research flowchart** — the paper's goal, data, methodology, and findings as a visual diagram you can absorb in seconds. - A **complexity vs. rigor score** — how mathematically demanding the paper is, and how empirically tested its claims are. See the [scoring methodology](/score-guide/). - A **plain-language summary** with extracted keywords. ## Why scores? Most paper triage fails in one of two ways: you either drown in elegant theory that has never touched real data, or you chase backtests built on shaky foundations. Scoring every paper on **math complexity** and **empirical rigor** separately makes that trade-off visible before you invest an afternoon in reading. The four quadrants — Holy Grail, Street Traders, Lab Rats, and Philosophers — are explained in the [score guide](/score-guide/). ## How it works The pipeline fetches new papers daily, extracts the abstract and structure, scores each paper with an LLM-assisted review against a fixed rubric, and renders the flowchart. Scores are heuristics, not peer review — they tell you where to look first, not what to believe. ## Contact Questions, corrections, or partnership inquiries: reach out on [LinkedIn](https://www.linkedin.com/in/aminehadbi/) or subscribe to the newsletter below — replies go straight to the author. ---------------------------------------------------------------------- # A Checklist for Reproducing Quant Research URL: https://thequant.space/guides/reproduce-quant-research/ Section: Guides Date: 2026-09-07 Description: A six-stage checklist for reproducing quantitative finance papers: acquisition, alignment, independent reimplementation, reconciliation, stress testing, and documentation — with the failure modes at each stage. A paper you've read is a story; a paper you've reproduced is a result. The gap between the two is where most published alpha lives — and dies. This is the checklist we use when a paper [survives triage](/guides/how-to-read-quant-finance-papers/) and [the backtest looks real](/guides/evaluate-trading-backtest/): six stages, each with its characteristic trap. A condensed version is downloadable at the end. ## Stage 1 — Acquire: establish what exists - ☐ **Paper version**: get the latest revision (arXiv papers mutate; results sometimes soften between v1 and v4 — diff the tables). - ☐ **Code**: linked repo? Note the commit hash and whether it actually produces the paper's tables or is "illustrative." Most repos are the latter. - ☐ **Data**: can you buy or download the *same* data — same vendor, frequency, adjustments? "CRSP 1963–2018" is reproducible; "proprietary tick data" ends the exercise honestly (mark it unreproducible, not wrong). - ☐ **Compute budget**: estimate before starting. A 5-minute daily-bars study and a GPU-month of deep-model training are different commitments — [price it first](/tools/compute-cost-calculator/). **The trap:** starting the reimplementation before confirming the data is obtainable. Data access decides reproducibility more often than method complexity does — which is why [vendor choice](/guides/market-data-vendors/) is a research-methodology issue, not procurement. ## Stage 2 — Align: match their world before testing it - ☐ Universe construction rules (exchanges, price floors, liquidity filters, share classes) reproduced *as of each date* — [survivorship-clean](/guides/survivorship-bias/). - ☐ Sample dates exactly; note anything the window conveniently excludes (2008? 2020? 2022?). - ☐ Definitions pinned: "monthly return" (close-to-close? which day?), "volatility" (realized? which estimator?), "momentum" (12-1? 6-1?). Every undefined term is a fork in the road. - ☐ Adjustment policy: splits, dividends, [restatements](/guides/look-ahead-bias-point-in-time-data/) — matched to what the paper's vendor would have shown *at the time*. **The trap:** silently using your own conventions. When your numbers diverge later, you won't know if the paper is wrong or your alignment is. ## Stage 3 — Reimplement: independently, on purpose - ☐ Write the pipeline from the paper's description *before* reading their code. Their code is the answer key, and answer keys teach nothing when copied — worse, **porting their code ports their bugs**. (Our [T-KAN replication review](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion) found the headline comparison rested on a baseline whose LSTM was never trained — invisible unless you reimplement or audit, not copy.) - ☐ One reproducible environment: pinned dependencies, fixed seeds, logged config — the same discipline as [taking a strategy live](/guides/notebook-to-production/#gate-2--code). - ☐ Where the paper is ambiguous, log the decision and both interpretations. The ambiguity list *is* a research artifact. ## Stage 4 — Reconcile: match statistics before results The order matters enormously and everyone gets it backwards: - ☐ First match **descriptive statistics**: universe size per year, mean/vol of returns, coverage. If your universe has 2,900 stocks and theirs 3,400, stop — nothing downstream is interpretable. - ☐ Then match **intermediate artifacts**: signal distributions, factor loadings, portfolio turnover. - ☐ Only then compare **headline results** — with tolerance bands agreed in advance (sign and rough magnitude are success; the third decimal is not). - ☐ **Feed the pipeline a case with a known answer** and demand machine precision — the strongest validation pattern we know, per [the Monte Carlo audit](/guides/auditing-a-monte-carlo/#61-test-invariants-not-outputs). **The trap:** jumping straight to the Sharpe ratio. When it doesn't match, you now have one uninterpretable number and no idea which of forty upstream choices caused it. ## Stage 5 — Stress: the part the paper didn't do Reproduction confirms the claim *as stated*; stress testing asks whether it matters: - ☐ Costs at 1×/2×/3× ([the sensitivity curve](/guides/evaluate-trading-backtest/#3-what-happens-at-2-the-assumed-costs)). - ☐ Subperiods and regimes — does it exist outside one lucky era? - ☐ Parameter neighborhood — plateau or spike? - ☐ Seeds and data perturbations for ML papers ([multiple runs, not one](/topics/machine-learning/)). - ☐ Post-publication sample: the cleanest out-of-sample there is, and the reason [reproducing older papers is more informative](/guides/statistical-vs-economic-significance/) than reproducing last month's. ## Stage 6 — Document: make your reproduction reproducible - ☐ A short report: what matched, what diverged, the ambiguity list, the stress results. - ☐ Publish it if you can — replication notes are among the most valuable, least supplied artifacts in the field (and exactly what the discussion sections on our [paper pages](/flowcharts/) exist for). - ☐ Archive data snapshots and environment — per the [raw-zone discipline](/guides/quant-research-stack/#layer-1-market-data--own-your-pipeline-rent-the-feed). ## The honest outcome distribution Expect roughly: a third of attempts reproduce cleanly, a third reproduce weaker (the median published effect shrinks out of the author's hands), and a third fail on data access, ambiguity, or genuine error. All three outcomes are wins — the failed reproduction that costs you two weeks is the strategy that would have cost you a drawdown. **📥 Take it with you:** download the condensed checklist (.md) — drop it into the repo of your next reproduction attempt. --- **Companion frameworks:** [Is the backtest real? →](/guides/evaluate-trading-backtest/) · [Auditing a Monte Carlo →](/guides/auditing-a-monte-carlo/) ---------------------------------------------------------------------- # A Minimal ML Experiment-Tracking Stack for Quant Research URL: https://thequant.space/guides/ml-experiment-tracking/ Section: Guides Date: 2026-09-07 Description: Experiment tracking for quant ML: what to record, the finance-specific requirements (trial counts, temporal splits, leakage audits), tool tiers from SQLite to MLflow/W&B, and the minimal stack that suffices. Experiment tracking in mainstream ML answers "which config was best?" In quant research it must answer a second, harder question: **"how many things did we try?"** — because [the trial count is a statistical input](/guides/deflated-sharpe-ratio/#using-it-honestly-without-ceremony), and an untracked experiment is both irreproducible *and* an uncounted draw against your [luck budget](/guides/backtest-overfitting/). That dual purpose shapes everything below. ## What a run record must contain The [versioning guide's run row](/guides/versioning-datasets-backtests/#versioning-backtests-the-run-record) is the base: code commit, resolved-config hash, data-manifest versions, seed, [environment digest](/guides/docker-reproducible-research/), metrics, artifacts. Quant ML adds fields generic tools don't prompt for: - **Split protocol, recorded as data**: the exact [walk-forward/purge/embargo scheme](/guides/walk-forward-out-of-sample-testing/#the-walk-forward-structure) — fold boundaries and embargo widths as stored values, so "was this leaked?" is a query, not archaeology. - **Holdout-touch ledger**: which runs read the final test set. The [evaluate-once rule](/guides/walk-forward-out-of-sample-testing/#the-evaluate-once-rule) is only enforceable if touches are logged. - **Experiment lineage / family**: which [idea-log entry](/guides/quant-research-pipeline/#1-idea-intake--the-experiment-log) a run serves, and its parent run — the structure that makes N countable *per claim*, which is what [DSR](/guides/deflated-sharpe-ratio/) and [CSCV](/guides/cscv-explained/#honest-limitations) actually consume. - **Economic metrics beside ML metrics**: accuracy *and* [costed P&L translation](/guides/evaluate-ml-trading-paper/#the-worked-example-the-accuracy-illusion), turnover, per-fold Sharpe — recorded per run so the [accuracy-vs-P&L divergence](/guides/evaluate-ml-trading-paper/) is visible in the tracker, not discovered later. - **Seed batch, not seed**: [multi-seed dispersion](/guides/evaluate-rl-trading-paper/#the-rl-paper-checklist) as a first-class result; a run family with one seed is flagged incomplete. ## The minimal stack **Tier 0 — one table, one decorator** (where most solo shops should stay): a `runs` table in [your existing Postgres](/guides/parquet-vs-database/#the-workload--archetype-map), a `@tracked` decorator that snapshots config/commit/data-versions before and metrics after, artifacts to [dated directories](/guides/store-tick-data-efficiently/#the-reference-layout). ~150 lines, zero services, queryable with SQL — and because it's *your* schema, the finance fields above are just columns. This is the [experiment log](/guides/quant-research-pipeline/#1-idea-intake--the-experiment-log) and run registry unified. **Tier 1 — MLflow (self-hosted)**: adds a UI, artifact management, and model registry for the cost of [one more container](/guides/notebook-to-production/#1-one-vps-containers-as-the-unit-of-operation). Adopt when run volume makes SQL browsing tedious or a second person arrives. Keep the finance fields as tags/params — the schema discipline stays yours. **Tier 2 — W&B/hosted**: collaboration, sweeps, dashboards — real value at team scale; at solo scale mostly [a subscription and a data-residency question](/guides/quant-research-stack/#what-it-costs-september-2026-order-of-magnitude) (strategy configs are IP; check terms before they live on someone's cloud). The anti-pattern at every tier: **tracking as UI, not as constraint**. The tracker earns its keep when the [backtest harness *refuses* untracked runs](/guides/quant-research-pipeline/#4-backtest--a-harness-not-a-script) — the enforcement, not the dashboard, is the product. ## Sweeps: where tracking meets statistics Hyperparameter sweeps are industrialized [trial generation](/guides/backtest-overfitting/#how-it-happens-in-practice-ranked-by-frequency), so the sweep config itself is a research artifact: record the *grid searched* (not just winners), sweep-level N, and selection rule — then report the winner [deflated by that N](/guides/deflated-sharpe-ratio/), with the [parameter-plateau view](/guides/robustness-trading-strategy/#1-parameter-robustness--the-plateau-test) saved as a sweep artifact. A sweep whose losers were deleted is a [p-hacking machine with good UX](/guides/p-hacking-financial-research/#the-seven-practices-in-ascending-order-of-self-deception). ## Common tracking failures - **Notebook runs off the books**: the exploratory phase generates most trials and zero rows — [the moment a number matters, it goes through the harness](/guides/versioning-datasets-backtests/#common-versioning-failures). - **Deleted failures**: cleaning "clutter" destroys the denominator; failed runs are load-bearing data. - **Metrics without splits**: a tracked Sharpe with an untracked split protocol can't be audited for [leakage](/guides/look-ahead-bias-point-in-time-data/#6-normalization-and-statistics-computed-on-the-full-sample). - **Tracker-as-optional**: two paths through the codebase (tracked and quick) — the quick one wins, and the log lies. ## Questions to ask (of your setup, or a paper's) 1. Can you compute an honest N per headline claim, from records? 2. Are holdout touches logged and countable? 3. Do failed and abandoned runs persist? 4. Could a colleague rerun the best experiment [from its row alone](/guides/versioning-datasets-backtests/#the-lineage-chain-end-to-end)? ## What our scoring captures Papers rarely show trackers, but their shadows — disclosed search spaces, per-seed dispersion, split protocols stated as data — are among the [rigor axis's](/score-guide/) strongest ML signals, and their absence is why [the ML hub](/topics/machine-learning/) pairs high volume with high skepticism. Internally: this stack is what makes your own claims *scoreable* by your own standards. --- **The statistics it feeds:** [deflated Sharpe →](/guides/deflated-sharpe-ratio/) · [CSCV →](/guides/cscv-explained/) · **The pipeline home:** [research pipeline →](/guides/quant-research-pipeline/) · **The compute:** [GPU decisions →](/guides/local-vs-cloud-gpu/) ---------------------------------------------------------------------- # A Practical Quant Research Stack for a One-Person Shop (2026) URL: https://thequant.space/guides/quant-research-stack/ Section: Guides Date: 2026-09-06 Description: The complete toolchain for solo quant research in 2026: data, storage, backtesting, compute, and deployment — with honest costs and the mistakes to skip. Most solo quants assemble their stack backwards: they start with a backtesting framework, discover its data model fights their data vendor, bolt on a database that fits neither, and end up maintaining glue code instead of researching. This guide lays out the stack in dependency order — data first, alpha last — with what each layer costs and where spending money actually buys you research velocity. This is the setup we'd build today for **systematic research on daily-to-intraday horizons** at individual scale. High-frequency market making has different answers at every layer. ## The stack at a glance | Layer | Boring default | When to upgrade | |---|---|---| | Market data | One good API vendor (equities/crypto per your market) | Point-in-time fundamentals; tick history | | Storage | Parquet files + DuckDB | ClickHouse/TimescaleDB when data outgrows one machine | | Research | Python, Jupyter, pandas/Polars | Polars or Rust extensions when pandas becomes the bottleneck | | Backtesting | vectorbt or a ~500-line engine you wrote and understand | Event-driven engine only when strategy logic demands it | | Compute | Your workstation | One cloud GPU box, rented hourly, for ML experiments | | Execution | Broker API (IBKR-style) + a VPS | Colocation is a business decision, not a research one | | Monitoring | Cron + logs + one alerting channel | Grafana when you run more than two live strategies | ## Layer 1: Market data — own your pipeline, rent the feed The single highest-leverage decision in the stack. Three rules that survive contact with reality: **Store the raw vendor payloads.** Disk is cheap; re-downloading history after a vendor drops an endpoint, changes a symbology, or you churn off their plan is somewhere between expensive and impossible. Every API response lands in an append-only raw zone (compressed JSON or Parquet) before any transformation touches it. **Split prices from fundamentals mentally and financially.** End-of-day and intraday **price** data is commoditized — API vendors in the $30–$250/month bracket cover it well for US equities and crypto. **Point-in-time fundamentals** (as-reported, with restatement history) are not commoditized, and cheap fundamental feeds silently backfill restated numbers — a look-ahead bias machine. If your strategies use fundamentals, this is where the real data budget goes; if not, don't pay for them. **Survivorship is the bias that kills retail backtests.** Whatever vendor you choose, confirm delisted securities are in the historical universe, and test it: query a ticker that went bankrupt and check the data is there. Our [market data vendor guide](/guides/market-data-vendors/) goes deep on this. ## Layer 2: Storage — Parquet + DuckDB until it hurts The 2026 answer for a one-person shop is boring and excellent: **partitioned Parquet files on local NVMe, queried with DuckDB**. - Parquet is vendor-neutral: every tool in the ecosystem reads it, so no migration ever strands your data. - DuckDB runs full analytical SQL over those files in-process — no server, no ops, and on a modern workstation it scans years of daily bars in milliseconds and single-name tick days in seconds. - Partition by date (and by symbol only if you routinely query single names); compress with zstd; keep a manifest of what's loaded. You upgrade to a real database server when one of these becomes true: multiple processes need concurrent writes, the working set outgrows one machine, or you're running continuous intraday ingestion. At that point the candidates are ClickHouse, TimescaleDB, QuestDB, or kdb+ — compared honestly in our [tick database guide](/guides/tick-data-databases/). ## Layer 3: Research environment — optimize for iteration speed Python remains the only defensible default. The 2026 refinements: - **Polars over pandas for anything heavy.** The API is stricter, the speedups on groupbys and joins are routinely 5–20× on research workloads, and it shares Arrow memory with DuckDB so data moves between them without copies. - **One repository, one environment.** A single `uv`-managed project with your data loaders, feature library, and notebooks beats a constellation of per-idea folders. Your feature definitions are your real IP; version them like production code. - **Notebooks are for exploration, modules are for truth.** The moment a computation matters, it moves from notebook to a tested function. Every serious research shop converges on this rule after their first irreproducible result — skip the learning fee. ## Layer 4: Backtesting — the layer where honesty lives Framework choice matters less than the discipline around it, but the practical options in 2026: - **vectorbt** for fast vectorized scans over signal grids — ideal for the "is there anything here at all?" phase. - **A small event-driven engine you wrote yourself** (a few hundred lines) for anything approaching production. Not because existing engines are bad, but because every assumption you didn't write is an assumption you can't audit — and fills, costs, and timing assumptions are where backtests lie. - Whatever you use, hold the line on the big four: **point-in-time data, explicit transaction costs, no same-bar execution on signal bars, and a untouched out-of-sample period**. The papers that score highest on empirical rigor across our [topic hubs](/topics/) share exactly these habits; so should you. ## Layer 5: Compute — rent spikes, own the baseline A capable workstation (fast NVMe, 64–128GB RAM, a mid-range GPU) covers 90% of systematic research. For the ML-heavy remainder — hyperparameter sweeps, deep models on intraday data — **hourly cloud GPUs** (the spot/community tier of the GPU cloud market) are dramatically cheaper than owning hardware you use in bursts. The workflow that keeps costs sane: develop and debug locally on subsampled data, rent for full-scale runs, pull artifacts back, kill the instance. If your monthly rental bill persistently exceeds the amortized cost of a used workstation GPU, buy the GPU. ## Layer 6: Deployment — a $10 VPS and a kill switch For daily-to-hourly strategies: a small VPS in a reliable region, your strategy as a systemd service or container, orders through your broker's API, and — non-negotiable — an independent **kill switch**: a separate process that flattens positions if the strategy stops heartbeating or drawdown breaches a hard limit. Latency optimization below the hundreds-of-milliseconds level is wasted effort at these horizons; reliability engineering is not. ## What it costs (September 2026, order-of-magnitude) - **Floor (~$50/month):** one data API, local storage, your existing computer, free-tier monitoring. Fully sufficient for daily-frequency equities/crypto research. - **Comfortable (~$200–500/month):** better data (intraday history, a fundamentals feed), a VPS, occasional GPU rental. - **The trap:** paying for infrastructure before a strategy earns it. Every layer above the floor should be purchased *in response to a specific bottleneck*, not in anticipation of one. *Prices and tool choices dated September 2026; verify before committing. This guide contains no sponsored placements — if that ever changes, it will be disclosed inline.* --- **Next:** [How to choose a market data vendor →](/guides/market-data-vendors/) ---------------------------------------------------------------------- # A Visual Map of Alternative-Data Research URL: https://thequant.space/guides/map-alternative-data/ Section: Guides Date: 2026-09-07 Description: A Visual Map of Alternative-Data Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored. Alternative-data research asks what non-price information is worth, and its history is a decay chronicle: each source (news tone, tweets, search volume) earns until it's crowded. The current wave is LLM-extracted signals, which adds a new failure mode — [training-cutoff leakage](/guides/look-ahead-bias-point-in-time-data/#7-llm-training-data-leakage--the-newest-member) — to the old ones. Every branch below should be read with the decay clock in mind: publication date is a feature, not metadata. ```mermaid flowchart TD ROOT["Alternative-Data Research
805 papers"] ROOT --> S0["News & sentiment
243 papers"] ROOT --> S1["Social & attention data
32 papers"] ROOT --> S2["LLM-extracted signals
336 papers"] ROOT --> S3["Earnings calls & filings
10 papers"] ROOT --> S4["Novel sources
6 papers"] ``` *805 papers mapped · ranked by [rigor-weighted score](/score-guide/) · regenerated automatically as the [daily pipeline](/about/) scores new papers.* ## News & sentiment The original alternative data: tone, coverage, and reaction. - [Returns and Order Flow Imbalances: Intraday Dynamics and Macroeconomic News Effects](/flowcharts/returns-and-order-flow-imbalances-intraday-dynamics-and-mac/) — *Holy Grail*, rigor 9.0 - [NewsNet-SDF: Stochastic Discount Factor Estimation with Pretrained Language Model News Embeddings via Adversarial Networks](/flowcharts/newsnet-sdf-stochastic-discount-factor-estimation-with-pret/) — *Holy Grail*, rigor 9.0 - [HODL Strategy or Fantasy? 480 Million Crypto Market Simulations and the Macro-Sentiment Effect](/flowcharts/hodl-strategy-or-fantasy-480-million-crypto-market-simulati/) — *Holy Grail*, rigor 9.0 - [Transformer-based CoVaR: Systemic Risk in Textual Information](/flowcharts/transformer-based-covar-systemic-risk-in-textual-information/) — *Holy Grail*, rigor 8.5 - [Integrating Large Language Models and Reinforcement Learning for Sentiment-Driven Quantitative Trading](/flowcharts/integrating-large-language-models-and-reinforcement-learning/) — *Holy Grail*, rigor 8.5 - [Deep g-Pricing for CSI 300 Index Options with Volatility Trajectories and Market Sentiment](/flowcharts/deep-g-pricing-for-csi-300-index-options-with-volatility-trajectories-and/) — *Holy Grail*, rigor 8.5 - …and 237 more in the [news & sentiment area](/topics/nlp-llm/) ## Social & attention data Tweets, boards, and search as attention proxies. - [Quantum Adaptive Self-Attention for Financial Rebalancing: An Empirical Study on Automated Market Makers in Decentralized Finance](/flowcharts/quantum-adaptive-self-attention-for-financial-rebalancing-a/) — *Holy Grail*, rigor 9.0 - [Uni-FinLLM: A Unified Multimodal Large Language Model with Modular Task Heads for Micro-Level Stock Prediction and Macro-Level Systemic Risk Assessment](/flowcharts/uni-finllm-a-unified-multimodal-large-language-model-with-m/) — *Holy Grail*, rigor 8.5 - [Decision-informed Neural Networks with Large Language Model Integration for Portfolio Optimization](/flowcharts/decision-informed-neural-networks-with-large-language-model/) — *Holy Grail*, rigor 7.0 - [Economic uncertainty and exchange rates linkage revisited: modelling tail dependence with high frequency data](/flowcharts/economic-uncertainty-and-exchange-rates-linkage-revisited-m/) — *Holy Grail*, rigor 7.0 - [RiskLabs: Predicting Financial Risk Using Large Language Model based on Multimodal and Multi-Sources Data](/flowcharts/risklabs-predicting-financial-risk-using-large-language-mod/) — *Holy Grail*, rigor 7.0 - [Looking into informal currency markets as Limit Order Books: impact of market makers](/flowcharts/looking-into-informal-currency-markets-as-limit-order-books/) — *Street Traders*, rigor 8.5 - …and 26 more in the [social & attention data area](/topics/nlp-llm/) ## LLM-extracted signals Language models as feature extractors — the current wave. - [Kronos: A Foundation Model for the Language of Financial Markets](/flowcharts/kronos-a-foundation-model-for-the-language-of-financial-mar/) — *Holy Grail*, rigor 9.0 - [A Hybrid Architecture for Options Wheel Strategy Decisions: LLM-Generated Bayesian Networks for Transparent Trading](/flowcharts/a-hybrid-architecture-for-options-wheel-strategy-decisions/) — *Holy Grail*, rigor 9.0 - [The Cross-Section of Stock Returns and AI Exposure](/flowcharts/the-cross-section-of-stock-returns-and-ai-exposure/) — *Holy Grail*, rigor 9.0 - [Hybrid LLM and Higher-Order Quantum Approximate Optimization for CSA Collateral Management](/flowcharts/hybrid-llm-and-higher-order-quantum-approximate-optimization/) — *Holy Grail*, rigor 7.5 - [Target alignment, dilution and forecast selection when cross-sectional forecasts share a common target](/flowcharts/target-alignment-dilution-and-forecast-selection-when-cross-sectional-forecasts/) — *Holy Grail*, rigor 7.5 - [Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation](/flowcharts/predicting-liquidity-aware-bond-yields-using-causal-gans-and/) — *Holy Grail*, rigor 8.0 - …and 330 more in the [llm-extracted signals area](/topics/nlp-llm/) ## Earnings calls & filings Structured corporate text and what leaks through it. - [Explainable AI for Comprehensive Risk Assessment for Financial Reports: A Lightweight Hierarchical Transformer Network Approach](/flowcharts/explainable-ai-for-comprehensive-risk-assessment-for-financial-reports-a/) — *Street Traders*, rigor 8.5 - [Generative AI, Managerial Expectations, and Economic Activity](/flowcharts/generative-ai-managerial-expectations-and-economic-activit/) — *Street Traders*, rigor 8.0 - [AMA-LSTM: Pioneering Robust and Fair Financial Audio Analysis for Stock Volatility Prediction](/flowcharts/ama-lstm-pioneering-robust-and-fair-financial-audio-analysi/) — *Holy Grail*, rigor 7.0 - [The Strategic Gap: How AI-Driven Timing and Complexity Shape Investor Trust in the Age of Digital Agents](/flowcharts/the-strategic-gap-how-ai-driven-timing-and-complexity-shape-investor-trust-in/) — *Street Traders*, rigor 7.5 - [Quantifying A Firm's AI Engagement: Constructing Objective, Data-Driven, AI Stock Indices Using 10-K Filings](/flowcharts/quantifying-a-firms-ai-engagement-constructing-objective/) — *Street Traders*, rigor 7.5 - [Cyber risk and the cross-section of stock returns](/flowcharts/cyber-risk-and-the-cross-section-of-stock-returns/) — *Street Traders*, rigor 7.0 - …and 4 more in the [earnings calls & filings area](/topics/nlp-llm/) ## Novel sources Satellites, transactions, and everything stranger. - [CaT-GNN: Enhancing Credit Card Fraud Detection via Causal Temporal Graph Neural Networks](/flowcharts/cat-gnn-enhancing-credit-card-fraud-detection-via-causal-te/) — *Holy Grail*, rigor 8.0 - [Neural and Time-Series Approaches for Pricing Weather Derivatives: Performance and Regime Adaptation Using Satellite Data](/flowcharts/neural-and-time-series-approaches-for-pricing-weather-deriva/) — *Holy Grail*, rigor 7.0 - [Thailand Asset Value Estimation Using Aerial or Satellite Imagery](/flowcharts/thailand-asset-value-estimation-using-aerial-or-satellite-im/) — *Holy Grail*, rigor 7.0 - [TIMeSynC: Temporal Intent Modelling with Synchronized Context Encodings for Financial Service Applications](/flowcharts/timesync-temporal-intent-modelling-with-synchronized-context-encodings-for/) — *Street Traders*, rigor 6.5 - [Managing Basis Risks in Weather Parametric Insurance: A Quantitative Study of Diversification and Key Influencing Factors](/flowcharts/managing-basis-risks-in-weather-parametric-insurance-a-quantitative-study-of/) — *Street Traders*, rigor 5.0 - [Feasibility-First Satellite Integration in Robust Portfolio Architectures](/flowcharts/feasibility-first-satellite-integration-in-robust-portfolio/) — *Lab Rats*, rigor 2.0 --- **Evaluation lens for this field:** see the [guides index](/guides/) · **Full hub:** [/topics/nlp-llm/](/topics/nlp-llm/) ---------------------------------------------------------------------- # A Visual Map of Limit-Order-Book Prediction Research URL: https://thequant.space/guides/map-lob-prediction/ Section: Guides Date: 2026-09-07 Description: A Visual Map of Limit-Order-Book Prediction Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored. LOB prediction asks whether the visible book predicts the next move — and the honest answer is 'yes, weakly, at horizons where it's hard to monetize.' The literature runs from handcrafted imbalance features through the DeepLOB lineage to transformers and beyond, with evaluation pitfalls (label construction, [horizon overlap](/guides/walk-forward-out-of-sample-testing/), fee-blind backtests) recurring in every generation. Our own [T-KAN replication review](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion) is a live case study in branch four. ```mermaid flowchart TD ROOT["Limit-Order-Book Prediction Research
623 papers"] ROOT --> S0["Imbalance & handcrafted features
122 papers"] ROOT --> S1["Deep LOB models
47 papers"] ROOT --> S2["Transformers & new architectures
17 papers"] ROOT --> S3["Evaluation & pitfalls
56 papers"] ROOT --> S4["Price impact & flow dynamics
112 papers"] ``` *623 papers mapped · ranked by [rigor-weighted score](/score-guide/) · regenerated automatically as the [daily pipeline](/about/) scores new papers.* ## Imbalance & handcrafted features What the book's shape measures before any learning. - [Returns and Order Flow Imbalances: Intraday Dynamics and Macroeconomic News Effects](/flowcharts/returns-and-order-flow-imbalances-intraday-dynamics-and-mac/) — *Holy Grail*, rigor 9.0 - [Deep Learning Meets Queue-Reactive: A Framework for Realistic Limit Order Book Simulation](/flowcharts/deep-learning-meets-queue-reactive-a-framework-for-realisti/) — *Holy Grail*, rigor 9.0 - [The Limits of Complexity: Why Feature Engineering Beats Deep Learning in Investor Flow Prediction](/flowcharts/the-limits-of-complexity-why-feature-engineering-beats-deep/) — *Holy Grail*, rigor 9.0 - [Trade Execution Flow as the Underlying Source of Market Dynamics](/flowcharts/trade-execution-flow-as-the-underlying-source-of-market-dyna/) — *Holy Grail*, rigor 8.0 - [An Impulse Control Approach to Market Making in a Hawkes LOB Market](/flowcharts/an-impulse-control-approach-to-market-making-in-a-hawkes-lob/) — *Holy Grail*, rigor 7.8 - [The Subtle Interplay between Square-root Impact, Order Imbalance & Volatility: A Unifying Framework](/flowcharts/the-subtle-interplay-between-square-root-impact-order-imbalance-volatility-a/) — *Holy Grail*, rigor 8.5 - …and 116 more in the [imbalance & handcrafted features area](/topics/market-microstructure/) ## Deep LOB models CNN/LSTM architectures on raw book states — the DeepLOB lineage. - [Latent Continuum of Regimes in Limit Order Book Dynamics](/flowcharts/latent-continuum-of-regimes-in-limit-order-book-dynamics/) — *Holy Grail*, rigor 9.0 - [Temporal Kolmogorov-Arnold Networks (T-KAN) for High-Frequency Limit Order Book Forecasting: Efficiency, Interpretability, and Alpha Decay](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/) — *Holy Grail*, rigor 9.0 - [Painting the market: generative diffusion models for financial limit order book simulation and forecasting](/flowcharts/painting-the-market-generative-diffusion-models-for-financi/) — *Holy Grail*, rigor 9.0 - [Meta-Learning Neural Process for Implied Volatility Surfaces with SABR-induced Priors](/flowcharts/meta-learning-neural-process-for-implied-volatility-surfaces/) — *Holy Grail*, rigor 8.0 - [Deep Attentive Survival Analysis in Limit Order Books: Estimating Fill Probabilities with Convolutional-Transformers](/flowcharts/deep-attentive-survival-analysis-in-limit-order-books-estim/) — *Holy Grail*, rigor 8.0 - [Multimodal Stock Price Prediction: A Case Study of the Russian Securities Market](/flowcharts/multimodal-stock-price-prediction-a-case-study-of-the-russi/) — *Holy Grail*, rigor 9.0 - …and 41 more in the [deep lob models area](/topics/market-microstructure/) ## Transformers & new architectures Attention, KANs, and the current frontier. - [Quantum Adaptive Self-Attention for Financial Rebalancing: An Empirical Study on Automated Market Makers in Decentralized Finance](/flowcharts/quantum-adaptive-self-attention-for-financial-rebalancing-a/) — *Holy Grail*, rigor 9.0 - [IVE: Enhanced Probabilistic Forecasting of Intraday Volume Ratio with Transformers](/flowcharts/ive-enhanced-probabilistic-forecasting-of-intraday-volume-r/) — *Holy Grail*, rigor 8.5 - [TRADES: Generating Realistic Market Simulations with Diffusion Models](/flowcharts/trades-generating-realistic-market-simulations-with-diffusi/) — *Holy Grail*, rigor 7.0 - [Event History Over Scale: Compact Transformers for Low-Latency Limit Order Book Forecasting](/flowcharts/event-history-over-scale-compact-transformers-for-low-latency-limit-order-book/) — *Holy Grail*, rigor 8.0 - [Cognitive Load and Information Processing in Financial Markets: Theory and Evidence from Disclosure Complexity](/flowcharts/cognitive-load-and-information-processing-in-financial-markets-theory-and/) — *Holy Grail*, rigor 7.5 - [IMM: An Imitative Reinforcement Learning Approach with Predictive Representation Learning for Automatic Market Making](/flowcharts/imm-an-imitative-reinforcement-learning-approach-with-predi/) — *Holy Grail*, rigor 6.0 - …and 11 more in the [transformers & new architectures area](/topics/market-microstructure/) ## Evaluation & pitfalls Labels, horizons, costs, and why accuracy ≠ P&L. - [Quantum and Classical Machine Learning in Decentralized Finance: Comparative Evidence from Multi-Asset Backtesting of Automated Market Makers](/flowcharts/quantum-and-classical-machine-learning-in-decentralized-fina/) — *Holy Grail*, rigor 9.0 - [Maximizing Battery Storage Profits via High-Frequency Intraday Trading](/flowcharts/maximizing-battery-storage-profits-via-high-frequency-intrad/) — *Holy Grail*, rigor 9.0 - [Utility-Weighted Forecasting and Calibration for Risk-Adjusted Decisions under Trading Frictions](/flowcharts/utility-weighted-forecasting-and-calibration-for-risk-adjusted-decisions-under/) — *Holy Grail*, rigor 8.5 - [Nonparametric Estimation of Self- and Cross-Impact](/flowcharts/nonparametric-estimation-of-self--and-cross-impact/) — *Holy Grail*, rigor 8.5 - [Resolution-Aware Perpetual Futures on Binary Prediction Markets: An Empirical Risk-Design Framework Using Polymarket Data](/flowcharts/resolution-aware-perpetual-futures-on-binary-prediction-markets-an-empirical/) — *Holy Grail*, rigor 9.0 - [Limit Order Book Simulation and Trade Evaluation with $K$-Nearest-Neighbor Resampling](/flowcharts/limit-order-book-simulation-and-trade-evaluation-with-k-ne/) — *Holy Grail*, rigor 8.0 - …and 50 more in the [evaluation & pitfalls area](/topics/market-microstructure/) ## Price impact & flow dynamics The mechanism underneath: how flow moves price. - [Strict universality of the square-root law in price impact across stocks: a complete survey of the Tokyo stock exchange](/flowcharts/strict-universality-of-the-square-root-law-in-price-impact-a/) — *Holy Grail*, rigor 9.0 - [Prime Match: A Privacy-Preserving Inventory Matching System](/flowcharts/prime-match-a-privacy-preserving-inventory-matching-system/) — *Holy Grail*, rigor 8.0 - [The double square-root law: Evidence for the mechanical origin of market impact using Tokyo Stock Exchange data](/flowcharts/the-double-square-root-law-evidence-for-the-mechanical-or/) — *Holy Grail*, rigor 8.5 - [Equity auction dynamics: latent liquidity models with activity acceleration](/flowcharts/equity-auction-dynamics-latent-liquidity-models-with-activi/) — *Holy Grail*, rigor 7.5 - [Extended State-dependent Hawkes Process for Limit Order Books: Mathematical Foundation and the Reproduction of Volatility Signature Plots](/flowcharts/extended-state-dependent-hawkes-process-for-limit-order-books-mathematical/) — *Holy Grail*, rigor 7.5 - [Signature approach for pricing and hedging path-dependent options with frictions](/flowcharts/signature-approach-for-pricing-and-hedging-path-dependent-op/) — *Holy Grail*, rigor 7.0 - …and 106 more in the [price impact & flow dynamics area](/topics/market-microstructure/) --- **Evaluation lens for this field:** see the [guides index](/guides/) · **Full hub:** [/topics/market-microstructure/](/topics/market-microstructure/) ---------------------------------------------------------------------- # A Visual Map of Machine Learning in Asset Pricing URL: https://thequant.space/guides/map-ml-asset-pricing/ Section: Guides Date: 2026-09-07 Description: A Visual Map of Machine Learning in Asset Pricing: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored. Machine learning met asset pricing in earnest when tree ensembles and neural nets were set loose on the cross-section of returns, and the field has been arguing about what they found ever since. The branches: predicting the cross-section, compressing it into latent factors, mining text for pricing information, and the growing literature on whether any of it survives [honest validation](/guides/evaluate-ml-trading-paper/). Trial counts matter more here than anywhere — read with the [deflated-Sharpe lens](/guides/deflated-sharpe-ratio/). ```mermaid flowchart TD ROOT["Machine Learning in Asset Pricing
1742 papers"] ROOT --> S0["Cross-sectional return prediction
186 papers"] ROOT --> S1["Latent factors & autoencoders
213 papers"] ROOT --> S2["Deep architectures
653 papers"] ROOT --> S3["Text & NLP pricing signals
110 papers"] ROOT --> S4["Skepticism & validation
29 papers"] ``` *1742 papers mapped · ranked by [rigor-weighted score](/score-guide/) · regenerated automatically as the [daily pipeline](/about/) scores new papers.* ## Cross-sectional return prediction The Gu–Kelly–Xiu lineage: ML on characteristics. - [Robust MCVaR Portfolio Optimization with Ellipsoidal Support and Reproducing Kernel Hilbert Space-based Uncertainty](/flowcharts/robust-mcvar-portfolio-optimization-with-ellipsoidal-support/) — *Holy Grail*, rigor 9.0 - [Scaling Conditional Autoencoders for Portfolio Optimization via Uncertainty-Aware Factor Selection](/flowcharts/scaling-conditional-autoencoders-for-portfolio-optimization/) — *Holy Grail*, rigor 9.0 - [Machine Learning Enhanced Multi-Factor Quantitative Trading: A Cross-Sectional Portfolio Optimization Approach with Bias Correction](/flowcharts/machine-learning-enhanced-multi-factor-quantitative-trading/) — *Holy Grail*, rigor 9.0 - [StockGPT: A GenAI Model for Stock Prediction and Trading](/flowcharts/stockgpt-a-genai-model-for-stock-prediction-and-trading/) — *Holy Grail*, rigor 9.0 - [Large (and Deep) Factor Models](/flowcharts/large-and-deep-factor-models/) — *Holy Grail*, rigor 8.0 - [The Cross-Section of Stock Returns and AI Exposure](/flowcharts/the-cross-section-of-stock-returns-and-ai-exposure/) — *Holy Grail*, rigor 9.0 - …and 180 more in the [cross-sectional return prediction area](/topics/machine-learning/) ## Latent factors & autoencoders Learning the factor structure instead of assuming it. - [Deep Learning Meets Queue-Reactive: A Framework for Realistic Limit Order Book Simulation](/flowcharts/deep-learning-meets-queue-reactive-a-framework-for-realisti/) — *Holy Grail*, rigor 9.0 - [A multi-factor model for improved commodity pricing: Calibration and an application to the oil market](/flowcharts/a-multi-factor-model-for-improved-commodity-pricing-calibra/) — *Holy Grail*, rigor 8.5 - [A Spatio-Temporal Machine Learning Model for Mortgage Credit Risk: Default Probabilities and Loan Portfolios](/flowcharts/a-spatio-temporal-machine-learning-model-for-mortgage-credit/) — *Holy Grail*, rigor 8.5 - [Spiking Neural Network for Cross-Market Portfolio Optimization in Financial Markets: A Neuromorphic Computing Approach](/flowcharts/spiking-neural-network-for-cross-market-portfolio-optimizati/) — *Holy Grail*, rigor 8.0 - [Dynamic Factor Correlation Model](/flowcharts/dynamic-factor-correlation-model/) — *Holy Grail*, rigor 8.5 - [Machine Learning Based Stress Testing Framework for Indian Financial Market Portfolios](/flowcharts/machine-learning-based-stress-testing-framework-for-indian-f/) — *Holy Grail*, rigor 8.0 - …and 207 more in the [latent factors & autoencoders area](/topics/machine-learning/) ## Deep architectures Networks, transformers, and attention on pricing problems. - [Quantum Adaptive Self-Attention for Financial Rebalancing: An Empirical Study on Automated Market Makers in Decentralized Finance](/flowcharts/quantum-adaptive-self-attention-for-financial-rebalancing-a/) — *Holy Grail*, rigor 9.0 - [Advancing Algorithmic Trading: A Multi-Technique Enhancement of Deep Q-Network Models](/flowcharts/advancing-algorithmic-trading-a-multi-technique-enhancement/) — *Holy Grail*, rigor 9.0 - [Optimizing Portfolio with Two-Sided Transactions and Lending: A Reinforcement Learning Framework](/flowcharts/optimizing-portfolio-with-two-sided-transactions-and-lending/) — *Holy Grail*, rigor 9.0 - [Temporal Kolmogorov-Arnold Networks (T-KAN) for High-Frequency Limit Order Book Forecasting: Efficiency, Interpretability, and Alpha Decay](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/) — *Holy Grail*, rigor 9.0 - [Causal and Predictive Modeling of Short-Horizon Market Risk and Systematic Alpha Generation Using Hybrid Machine Learning Ensembles](/flowcharts/causal-and-predictive-modeling-of-short-horizon-market-risk/) — *Holy Grail*, rigor 9.0 - [Few-Shot Learning Patterns in Financial Time-Series for Trend-Following Strategies](/flowcharts/few-shot-learning-patterns-in-financial-time-series-for-tren/) — *Holy Grail*, rigor 9.0 - …and 647 more in the [deep architectures area](/topics/machine-learning/) ## Text & NLP pricing signals News, filings, and language models as pricing inputs. - [HODL Strategy or Fantasy? 480 Million Crypto Market Simulations and the Macro-Sentiment Effect](/flowcharts/hodl-strategy-or-fantasy-480-million-crypto-market-simulati/) — *Holy Grail*, rigor 9.0 - [What Teaches Robots to Walk, Teaches Them to Trade too -- Regime Adaptive Execution using Informed Data and LLMs](/flowcharts/what-teaches-robots-to-walk-teaches-them-to-trade-too----re/) — *Holy Grail*, rigor 8.0 - [FactorBench: A Portfolio-Aware Benchmark for Automated Factor Mining](/flowcharts/factorbench-a-portfolio-aware-benchmark-for-automated-factor-mining/) — *Holy Grail*, rigor 9.0 - [Clustered Network Connectedness: A New Measurement Framework with Application to Global Equity Markets](/flowcharts/clustered-network-connectedness-a-new-measurement-framework/) — *Holy Grail*, rigor 8.0 - [Mamba Outpaces Reformer in Stock Prediction with Sentiments from Top Ten LLMs](/flowcharts/mamba-outpaces-reformer-in-stock-prediction-with-sentiments/) — *Holy Grail*, rigor 8.0 - [Interpretable Systematic Risk around the Clock](/flowcharts/interpretable-systematic-risk-around-the-clock/) — *Holy Grail*, rigor 8.5 - …and 104 more in the [text & nlp pricing signals area](/topics/machine-learning/) ## Skepticism & validation Replication failures, overfitting critiques, and honest benchmarks. - [Predicting Market Troughs: A Machine Learning Approach with Causal Interpretation](/flowcharts/predicting-market-troughs-a-machine-learning-approach-with/) — *Holy Grail*, rigor 9.0 - [Robust Utility Optimization via a GAN Approach](/flowcharts/robust-utility-optimization-via-a-gan-approach/) — *Holy Grail*, rigor 7.5 - [Beyond Correlation: Positive Definite Dependence Measures for Robust Inference, Flexible Scenarios, and Causal Modeling for Financial Portfolios](/flowcharts/beyond-correlation-positive-definite-dependence-measures-fo/) — *Holy Grail*, rigor 7.5 - [Variable Clustering via Distributionally Robust Nodewise Regression](/flowcharts/variable-clustering-via-distributionally-robust-nodewise-regression/) — *Holy Grail*, rigor 7.5 - [Robust valuation and optimal harvesting of forestry resources in the presence of catastrophe risk and parameter uncertainty](/flowcharts/robust-valuation-and-optimal-harvesting-of-forestry-resources-in-the-presence/) — *Holy Grail*, rigor 7.5 - [Dynamically Consistent Analysis of Realized Covariations in Term Structure Models](/flowcharts/dynamically-consistent-analysis-of-realized-covariations-in/) — *Holy Grail*, rigor 7.0 - …and 23 more in the [skepticism & validation area](/topics/machine-learning/) --- **Evaluation lens for this field:** see the [guides index](/guides/) · **Full hub:** [/topics/machine-learning/](/topics/machine-learning/) ---------------------------------------------------------------------- # A Visual Map of Market-Making Research URL: https://thequant.space/guides/map-market-making/ Section: Guides Date: 2026-09-07 Description: A Visual Map of Market-Making Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored. Market-making research answers one question three ways: how should a liquidity provider set quotes given inventory risk and adverse selection? The stochastic-control lineage (Avellaneda–Stoikov and descendants) answers with closed forms; the RL lineage answers with learned policies; the DeFi lineage redesigns the mechanism itself. The branches below follow those lines — with the [venue-specificity caveat](/guides/evaluate-microstructure-paper/) that applies doubly to anything empirical here. ```mermaid flowchart TD ROOT["Market-Making Research
611 papers"] ROOT --> S0["Inventory & optimal quoting
83 papers"] ROOT --> S1["Adverse selection & information
42 papers"] ROOT --> S2["RL market making
77 papers"] ROOT --> S3["AMMs & DeFi liquidity
103 papers"] ROOT --> S4["Empirical market making
114 papers"] ``` *611 papers mapped · ranked by [rigor-weighted score](/score-guide/) · regenerated automatically as the [daily pipeline](/about/) scores new papers.* ## Inventory & optimal quoting The stochastic-control core: quotes as a function of inventory and time. - [Risk-Sensitive Option Market Making with Arbitrage-Free eSSVI Surfaces: A Constrained RL and Stochastic Control Bridge](/flowcharts/risk-sensitive-option-market-making-with-arbitrage-free-essv/) — *Holy Grail*, rigor 7.8 - [An Impulse Control Approach to Market Making in a Hawkes LOB Market](/flowcharts/an-impulse-control-approach-to-market-making-in-a-hawkes-lob/) — *Holy Grail*, rigor 7.8 - [Prime Match: A Privacy-Preserving Inventory Matching System](/flowcharts/prime-match-a-privacy-preserving-inventory-matching-system/) — *Holy Grail*, rigor 8.0 - [Optimal Rebalancing in Dynamic AMMs](/flowcharts/optimal-rebalancing-in-dynamic-amms/) — *Holy Grail*, rigor 8.0 - [Signature approach for pricing and hedging path-dependent options with frictions](/flowcharts/signature-approach-for-pricing-and-hedging-path-dependent-op/) — *Holy Grail*, rigor 7.0 - [Adaptive Optimal Market Making Strategies with Inventory Liquidation Cos](/flowcharts/adaptive-optimal-market-making-strategies-with-inventory-liq/) — *Holy Grail*, rigor 8.0 - …and 77 more in the [inventory & optimal quoting area](/topics/market-microstructure/) ## Adverse selection & information Who picks you off, and what spreads must charge for it. - [ForesightFlow: An Information Leakage Score Framework for Prediction Markets](/flowcharts/foresightflow-an-information-leakage-score-framework-for-prediction-markets/) — *Holy Grail*, rigor 9.0 - [Fragmentation and optimal liquidity supply on decentralized exchanges](/flowcharts/fragmentation-and-optimal-liquidity-supply-on-decentralized/) — *Holy Grail*, rigor 8.0 - [Optimal Signal Extraction from Order Flow: A Matched Filter Perspective on Normalization and Market Microstructure](/flowcharts/optimal-signal-extraction-from-order-flow-a-matched-filter/) — *Holy Grail*, rigor 8.0 - [The Value of Information: A Puzzle](/flowcharts/the-value-of-information-a-puzzle/) — *Holy Grail*, rigor 8.0 - [Market Simulation under Adverse Selection](/flowcharts/market-simulation-under-adverse-selection/) — *Holy Grail*, rigor 7.0 - [Not All LPs Are Equal: The Active-Passive Gap in Automated Market Maker Liquidity Provision](/flowcharts/not-all-lps-are-equal-the-active-passive-gap-in-automated-market-maker/) — *Holy Grail*, rigor 8.0 - …and 36 more in the [adverse selection & information area](/topics/market-microstructure/) ## RL market making Learned quoting policies and their evaluation problems. - [Minimal Batch Adaptive Learning Policy Engine for Real-Time Mid-Price Forecasting in High-Frequency Trading](/flowcharts/minimal-batch-adaptive-learning-policy-engine-for-real-time/) — *Holy Grail*, rigor 8.5 - [Interpretable Hypothesis-Driven Trading:A Rigorous Walk-Forward Validation Framework for Market Microstructure Signals](/flowcharts/interpretable-hypothesis-driven-tradinga-rigorous-walk-forw/) — *Holy Grail*, rigor 9.0 - [Reinforcement Learning in Queue-Reactive Models: Application to Optimal Execution](/flowcharts/reinforcement-learning-in-queue-reactive-models-application/) — *Holy Grail*, rigor 8.0 - [RAmmStein: Regime Adaptation in Mean-reverting Markets with Stein Thresholds -- Optimal Impulse Control in Concentrated AMMs](/flowcharts/rammstein-regime-adaptation-in-mean-reverting-markets-with-stein-thresholds/) — *Holy Grail*, rigor 8.5 - [Adaptive Dueling Double Deep Q-networks in Uniswap V3 Replication and Extension with Mamba](/flowcharts/adaptive-dueling-double-deep-q-networks-in-uniswap-v3-replic/) — *Holy Grail*, rigor 8.0 - [Event-Based Limit Order Book Simulation under a Neural Hawkes Process: Application in Market-Making](/flowcharts/event-based-limit-order-book-simulation-under-a-neural-hawke/) — *Holy Grail*, rigor 7.0 - …and 71 more in the [rl market making area](/topics/market-microstructure/) ## AMMs & DeFi liquidity Constant-function markets, LP returns, and impermanent loss. - [Quantum Adaptive Self-Attention for Financial Rebalancing: An Empirical Study on Automated Market Makers in Decentralized Finance](/flowcharts/quantum-adaptive-self-attention-for-financial-rebalancing-a/) — *Holy Grail*, rigor 9.0 - [Quantum and Classical Machine Learning in Decentralized Finance: Comparative Evidence from Multi-Asset Backtesting of Automated Market Makers](/flowcharts/quantum-and-classical-machine-learning-in-decentralized-fina/) — *Holy Grail*, rigor 9.0 - [Automated Market Making and Decentralized Finance](/flowcharts/automated-market-making-and-decentralized-finance/) — *Holy Grail*, rigor 8.0 - [Approaching multifractal complexity in decentralized cryptocurrency trading](/flowcharts/approaching-multifractal-complexity-in-decentralized-cryptoc/) — *Holy Grail*, rigor 8.0 - [Multiblock MEV opportunities & protections in dynamic AMMs](/flowcharts/multiblock-mev-opportunities--protections-in-dynamic-amms/) — *Holy Grail*, rigor 8.0 - [Option Pricing on Automated Market Maker Tokens](/flowcharts/option-pricing-on-automated-market-maker-tokens/) — *Holy Grail*, rigor 8.0 - …and 97 more in the [amms & defi liquidity area](/topics/market-microstructure/) ## Empirical market making Measured spreads, dealer behavior, and liquidity supply in real books. - [Latent Continuum of Regimes in Limit Order Book Dynamics](/flowcharts/latent-continuum-of-regimes-in-limit-order-book-dynamics/) — *Holy Grail*, rigor 9.0 - [Kronos: A Foundation Model for the Language of Financial Markets](/flowcharts/kronos-a-foundation-model-for-the-language-of-financial-mar/) — *Holy Grail*, rigor 9.0 - [Returns and Order Flow Imbalances: Intraday Dynamics and Macroeconomic News Effects](/flowcharts/returns-and-order-flow-imbalances-intraday-dynamics-and-mac/) — *Holy Grail*, rigor 9.0 - [Temporal Kolmogorov-Arnold Networks (T-KAN) for High-Frequency Limit Order Book Forecasting: Efficiency, Interpretability, and Alpha Decay](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/) — *Holy Grail*, rigor 9.0 - [Painting the market: generative diffusion models for financial limit order book simulation and forecasting](/flowcharts/painting-the-market-generative-diffusion-models-for-financi/) — *Holy Grail*, rigor 9.0 - [Strict universality of the square-root law in price impact across stocks: a complete survey of the Tokyo stock exchange](/flowcharts/strict-universality-of-the-square-root-law-in-price-impact-a/) — *Holy Grail*, rigor 9.0 - …and 108 more in the [empirical market making area](/topics/market-microstructure/) --- **Evaluation lens for this field:** see the [guides index](/guides/) · **Full hub:** [/topics/market-microstructure/](/topics/market-microstructure/) ---------------------------------------------------------------------- # A Visual Map of Reinforcement Learning for Trading URL: https://thequant.space/guides/map-rl-trading/ Section: Guides Date: 2026-09-07 Description: A Visual Map of Reinforcement Learning for Trading: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored. RL-for-trading research divides by problem structure, and the division predicts credibility: execution and hedging (dense feedback, clear costs) produce results that replicate; end-to-end portfolio alpha (sparse feedback, non-stationary environment) produces the field's most spectacular simulator victories. The map below follows that gradient — read every branch with the [environment-audit checklist](/guides/evaluate-rl-trading-paper/). ```mermaid flowchart TD ROOT["Reinforcement Learning for Trading
319 papers"] ROOT --> S0["Portfolio & allocation RL
142 papers"] ROOT --> S1["Execution RL
29 papers"] ROOT --> S2["Market-making RL
14 papers"] ROOT --> S3["Hedging RL
33 papers"] ROOT --> S4["Methods & training
57 papers"] ``` *319 papers mapped · ranked by [rigor-weighted score](/score-guide/) · regenerated automatically as the [daily pipeline](/about/) scores new papers.* ## Portfolio & allocation RL End-to-end policies from prices to weights — the ambitious branch. - [Advancing Algorithmic Trading: A Multi-Technique Enhancement of Deep Q-Network Models](/flowcharts/advancing-algorithmic-trading-a-multi-technique-enhancement/) — *Holy Grail*, rigor 9.0 - [Optimizing Portfolio with Two-Sided Transactions and Lending: A Reinforcement Learning Framework](/flowcharts/optimizing-portfolio-with-two-sided-transactions-and-lending/) — *Holy Grail*, rigor 9.0 - [AlphaPortfolio: Direct Construction Through Deep Reinforcement Learning and Interpretable AI](/flowcharts/alphaportfolio-direct-construction-through-deep-reinforceme/) — *Holy Grail*, rigor 9.0 - [FR-LUX: Friction-Aware, Regime-Conditioned Policy Optimization for Implementable Portfolio Management](/flowcharts/fr-lux-friction-aware-regime-conditioned-policy-optimizati/) — *Holy Grail*, rigor 8.0 - [CAD: Clustering And Deep Reinforcement Learning Based Multi-Period Portfolio Management Strategy](/flowcharts/cad-clustering-and-deep-reinforcement-learning-based-multi/) — *Holy Grail*, rigor 8.5 - [Distributionally Robust Deep Q-Learning](/flowcharts/distributionally-robust-deep-q-learning/) — *Holy Grail*, rigor 7.5 - …and 136 more in the [portfolio & allocation rl area](/topics/reinforcement-learning/) ## Execution RL Order placement and scheduling — the credible branch. - [Minimal Batch Adaptive Learning Policy Engine for Real-Time Mid-Price Forecasting in High-Frequency Trading](/flowcharts/minimal-batch-adaptive-learning-policy-engine-for-real-time/) — *Holy Grail*, rigor 8.5 - [Reinforcement Learning in Queue-Reactive Models: Application to Optimal Execution](/flowcharts/reinforcement-learning-in-queue-reactive-models-application/) — *Holy Grail*, rigor 8.0 - [Learning Market Making with Closing Auctions](/flowcharts/learning-market-making-with-closing-auctions/) — *Holy Grail*, rigor 8.0 - [Optimal Execution with Reinforcement Learning](/flowcharts/optimal-execution-with-reinforcement-learning/) — *Holy Grail*, rigor 8.0 - [What Teaches Robots to Walk, Teaches Them to Trade too -- Regime Adaptive Execution using Informed Data and LLMs](/flowcharts/what-teaches-robots-to-walk-teaches-them-to-trade-too----re/) — *Holy Grail*, rigor 8.0 - [Event-Based Limit Order Book Simulation under a Neural Hawkes Process: Application in Market-Making](/flowcharts/event-based-limit-order-book-simulation-under-a-neural-hawke/) — *Holy Grail*, rigor 7.0 - …and 23 more in the [execution rl area](/topics/reinforcement-learning/) ## Market-making RL Learned quoting under inventory and adverse selection. - [Deep Learning of Robust Market Making under Regime-Switching Order Flow](/flowcharts/deep-learning-of-robust-market-making-under-regime-switching-order-flow/) — *Holy Grail*, rigor 8.0 - [Improving DeFi Accessibility through Efficient Liquidity Provisioning with Deep Reinforcement Learning](/flowcharts/improving-defi-accessibility-through-efficient-liquidity-pro/) — *Holy Grail*, rigor 8.0 - [Reinforcement Learning in High-frequency Market Making](/flowcharts/reinforcement-learning-in-high-frequency-market-making/) — *Holy Grail*, rigor 6.0 - [ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility](/flowcharts/arl-based-multi-action-market-making-with-hawkes-processes-a/) — *Holy Grail*, rigor 6.5 - [Exploiting Risk-Aversion and Size-dependent fees in FX Trading with Fitted Natural Actor-Critic](/flowcharts/exploiting-risk-aversion-and-size-dependent-fees-in-fx-tradi/) — *Holy Grail*, rigor 6.0 - [Robust Market Making: To Quote, or not To Quote](/flowcharts/robust-market-making-to-quote-or-not-to-quote/) — *Holy Grail*, rigor 6.0 - …and 8 more in the [market-making rl area](/topics/reinforcement-learning/) ## Hedging RL Deep hedging under costs — the second credible branch. - [Application of Deep Reinforcement Learning to At-the-Money S&P 500 Options Hedging](/flowcharts/application-of-deep-reinforcement-learning-to-at-the-money-s/) — *Holy Grail*, rigor 7.5 - [Stochastic Policy Gradient Methods in the Uncertain Volatility Model](/flowcharts/stochastic-policy-gradient-methods-in-the-uncertain-volatility-model/) — *Holy Grail*, rigor 7.5 - [Deeper Hedging: A New Agent-based Model for Effective Deep Hedging](/flowcharts/deeper-hedging-a-new-agent-based-model-for-effective-deep-h/) — *Holy Grail*, rigor 7.5 - [Deep Hedging with Reinforcement Learning: A Practical Framework for Option Risk Management](/flowcharts/deep-hedging-with-reinforcement-learning-a-practical-framew/) — *Holy Grail*, rigor 8.5 - [Enhancing Deep Hedging of Options with Implied Volatility Surface Feedback Information](/flowcharts/enhancing-deep-hedging-of-options-with-implied-volatility-su/) — *Holy Grail*, rigor 7.5 - [A Risk Sensitive Contract-unified Reinforcement Learning Approach for Option Hedging](/flowcharts/a-risk-sensitive-contract-unified-reinforcement-learning-approach-for-option/) — *Holy Grail*, rigor 8.5 - …and 27 more in the [hedging rl area](/topics/reinforcement-learning/) ## Methods & training Algorithms, sample efficiency, and environment design. - [Integrating Large Language Models and Reinforcement Learning for Sentiment-Driven Quantitative Trading](/flowcharts/integrating-large-language-models-and-reinforcement-learning/) — *Holy Grail*, rigor 8.5 - [Interpretable Hypothesis-Driven Trading:A Rigorous Walk-Forward Validation Framework for Market Microstructure Signals](/flowcharts/interpretable-hypothesis-driven-tradinga-rigorous-walk-forw/) — *Holy Grail*, rigor 9.0 - [AlphaSAGE: Structure-Aware Alpha Mining via GFlowNets for Robust Exploration](/flowcharts/alphasage-structure-aware-alpha-mining-via-gflownets-for-ro/) — *Holy Grail*, rigor 8.0 - [Harnessing Deep Q-Learning for Enhanced Statistical Arbitrage in High-Frequency Trading: A Comprehensive Exploration](/flowcharts/harnessing-deep-q-learning-for-enhanced-statistical-arbitrag/) — *Holy Grail*, rigor 7.5 - [Bridging Econometrics and AI: VaR Estimation via Reinforcement Learning and GARCH Models](/flowcharts/bridging-econometrics-and-ai-var-estimation-via-reinforceme/) — *Holy Grail*, rigor 8.5 - [Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation](/flowcharts/predicting-liquidity-aware-bond-yields-using-causal-gans-and/) — *Holy Grail*, rigor 8.0 - …and 51 more in the [methods & training area](/topics/reinforcement-learning/) --- **Evaluation lens for this field:** see the [guides index](/guides/) · **Full hub:** [/topics/reinforcement-learning/](/topics/reinforcement-learning/) ---------------------------------------------------------------------- # A Visual Map of Statistical-Arbitrage Research URL: https://thequant.space/guides/map-statistical-arbitrage/ Section: Guides Date: 2026-09-07 Description: A Visual Map of Statistical-Arbitrage Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored. Statistical arbitrage research clusters into a few durable questions: how to find related instruments, how to model the spread, how to time entries, and how to survive the strategy's known decay. This map organizes the archive's stat-arb papers along those lines, ranked by our [rigor-weighted score](/score-guide/). Read it with the [stat-arb evaluation lens](/guides/evaluate-trading-backtest/): the classic distance method famously decayed post-2000s, so sample period is the first thing to check in every branch below. ```mermaid flowchart TD ROOT["Statistical-Arbitrage Research
90 papers"] ROOT --> S0["Pairs & cointegration
26 papers"] ROOT --> S1["Mean-reversion signals
42 papers"] ROOT --> S2["ML-driven stat arb
8 papers"] ROOT --> S3["Crypto & cross-venue
4 papers"] ROOT --> S4["Execution & capacity
2 papers"] ``` *90 papers mapped · ranked by [rigor-weighted score](/score-guide/) · regenerated automatically as the [daily pipeline](/about/) scores new papers.* ## Pairs & cointegration The classical core: finding and modeling co-moving pairs. - [Statistical arbitrage portfolio construction based on preference relations](/flowcharts/statistical-arbitrage-portfolio-construction-based-on-prefer/) — *Holy Grail*, rigor 8.5 - [Low Volatility Stock Portfolio Through High Dimensional Bayesian Cointegration](/flowcharts/low-volatility-stock-portfolio-through-high-dimensional-baye/) — *Holy Grail*, rigor 7.0 - [Statistical Arbitrage in Polish Equities Market Using Deep Learning Techniques](/flowcharts/statistical-arbitrage-in-polish-equities-market-using-deep-l/) — *Holy Grail*, rigor 8.5 - [Temporal Representation Learning for Stock Similarities and Its Applications in Investment Management](/flowcharts/temporal-representation-learning-for-stock-similarities-and/) — *Holy Grail*, rigor 8.0 - [Copula-Based Trading of Cointegrated Cryptocurrency Pairs](/flowcharts/copula-based-trading-of-cointegrated-cryptocurrency-pairs/) — *Holy Grail*, rigor 8.0 - [Price Discovery in Cryptocurrency Markets](/flowcharts/price-discovery-in-cryptocurrency-markets/) — *Holy Grail*, rigor 7.5 - …and 20 more in the [pairs & cointegration area](/topics/statistical-arbitrage/) ## Mean-reversion signals Spread construction and reversion timing beyond simple pairs. - [RAmmStein: Regime Adaptation in Mean-reverting Markets with Stein Thresholds -- Optimal Impulse Control in Concentrated AMMs](/flowcharts/rammstein-regime-adaptation-in-mean-reverting-markets-with-stein-thresholds/) — *Holy Grail*, rigor 8.5 - [Capturing Smile Dynamics with the Quintic Volatility Model: SPX, Skew-Stickiness Ratio and VIX](/flowcharts/capturing-smile-dynamics-with-the-quintic-volatility-model-spx-skew-stickiness/) — *Holy Grail*, rigor 8.5 - [Pricing energy spread options with variance gamma-driven Ornstein-Uhlenbeck dynamics](/flowcharts/pricing-energy-spread-options-with-variance-gamma-driven-ornstein-uhlenbeck/) — *Holy Grail*, rigor 7.5 - [A Mean-Reverting Model of Exchange Rate Risk Premium Using Ornstein-Uhlenbeck Dynamics](/flowcharts/a-mean-reverting-model-of-exchange-rate-risk-premium-using-o/) — *Holy Grail*, rigor 8.0 - [The quintic Ornstein-Uhlenbeck volatility model that jointly calibrates SPX & VIX smiles](/flowcharts/the-quintic-ornstein-uhlenbeck-volatility-model-that-jointly-calibrates-spx-vix/) — *Holy Grail*, rigor 7.5 - [Neural-Actuarial Longevity Forecasting: Anchoring LSTMs for Explainable Risk Management](/flowcharts/neural-actuarial-longevity-forecasting-anchoring-lstms-for-explainable-risk/) — *Holy Grail*, rigor 8.2 - …and 36 more in the [mean-reversion signals area](/topics/statistical-arbitrage/) ## ML-driven stat arb Learned similarity, embeddings, and nonlinear spread models. - [Statistical arbitrage in multi-pair trading strategy based on graph clustering algorithms in US equities market](/flowcharts/statistical-arbitrage-in-multi-pair-trading-strategy-based-o/) — *Holy Grail*, rigor 7.5 - [Harnessing Deep Q-Learning for Enhanced Statistical Arbitrage in High-Frequency Trading: A Comprehensive Exploration](/flowcharts/harnessing-deep-q-learning-for-enhanced-statistical-arbitrag/) — *Holy Grail*, rigor 7.5 - [Is the difference between deep hedging and delta hedging a statistical arbitrage?](/flowcharts/is-the-difference-between-deep-hedging-and-delta-hedging-a-s/) — *Holy Grail*, rigor 8.0 - [Machine Learning-based Relative Valuation of Municipal Bonds](/flowcharts/machine-learning-based-relative-valuation-of-municipal-bonds/) — *Holy Grail*, rigor 8.0 - [Hybrid Quantum-Classical Ensemble Learning for S&P 500 Directional Prediction](/flowcharts/hybrid-quantum-classical-ensemble-learning-for-s%5Cp-500-dire/) — *Holy Grail*, rigor 7.2 - [Reinforcement Learning Pair Trading: A Dynamic Scaling approach](/flowcharts/reinforcement-learning-pair-trading-a-dynamic-scaling-appro/) — *Holy Grail*, rigor 7.0 - …and 2 more in the [ml-driven stat arb area](/topics/statistical-arbitrage/) ## Crypto & cross-venue Stat arb where the venues are the anomaly. - [Automated Market Making and Decentralized Finance](/flowcharts/automated-market-making-and-decentralized-finance/) — *Holy Grail*, rigor 8.0 - [Graph Learning for Foreign Exchange Rate Prediction and Statistical Arbitrage](/flowcharts/graph-learning-for-foreign-exchange-rate-prediction-and-stat/) — *Holy Grail*, rigor 7.5 - [Decentralised Finance and Automated Market Making: Execution and Speculation](/flowcharts/decentralised-finance-and-automated-market-making-execution/) — *Holy Grail*, rigor 7.0 - [CTBench: Cryptocurrency Time Series Generation Benchmark](/flowcharts/ctbench-cryptocurrency-time-series-generation-benchmark/) — *Street Traders*, rigor 8.5 ## Execution & capacity Turning spreads into fills without eating the edge. - [On a fundamental statistical edge principle](/flowcharts/on-a-fundamental-statistical-edge-principle/) — *Lab Rats*, rigor 4.0 - [Pricing and Hedging Financial Derivatives in Merger\&Acquisition Deals with Price Impact](/flowcharts/pricing-and-hedging-financial-derivatives-in-merger-acquisition-deals-with/) — *Lab Rats*, rigor 3.0 --- **Evaluation lens for this field:** see the [guides index](/guides/) · **Full hub:** [/topics/statistical-arbitrage/](/topics/statistical-arbitrage/) ---------------------------------------------------------------------- # A Visual Map of Volatility Forecasting Research URL: https://thequant.space/guides/map-volatility-forecasting/ Section: Guides Date: 2026-09-07 Description: A Visual Map of Volatility Forecasting Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored. Volatility is the most forecastable object in finance, which makes its literature a rare pleasure: models compete on real out-of-sample benchmarks and the baseline (HAR-RV) is famously hard to beat. The branches run from the GARCH family through realized measures to implied-vol information and the deep-learning contenders — with the standing question for every new entrant: [does it beat HAR under an honest loss function](/guides/evaluate-trading-backtest/)? ```mermaid flowchart TD ROOT["Volatility Forecasting Research
746 papers"] ROOT --> S0["GARCH family & extensions
111 papers"] ROOT --> S1["Realized volatility & HAR
67 papers"] ROOT --> S2["Implied & options-based
248 papers"] ROOT --> S3["Rough & stochastic volatility
132 papers"] ROOT --> S4["ML volatility forecasting
41 papers"] ``` *746 papers mapped · ranked by [rigor-weighted score](/score-guide/) · regenerated automatically as the [daily pipeline](/about/) scores new papers.* ## GARCH family & extensions The parametric core and its many children. - [Dynamic allocation: extremes, tail dependence, and regime Shifts](/flowcharts/dynamic-allocation-extremes-tail-dependence-and-regime-sh/) — *Holy Grail*, rigor 9.0 - [Synthetic Financial Data Generation for Enhanced Financial Modelling](/flowcharts/synthetic-financial-data-generation-for-enhanced-financial-m/) — *Holy Grail*, rigor 9.0 - [Beyond the Mean: Limit Theory and Tests for Infinite-Mean Autoregressive Conditional Durations](/flowcharts/beyond-the-mean-limit-theory-and-tests-for-infinite-mean-au/) — *Holy Grail*, rigor 8.5 - [Neural Term Structure of Additive Process for Option Pricing](/flowcharts/neural-term-structure-of-additive-process-for-option-pricing/) — *Holy Grail*, rigor 8.0 - [Cluster GARCH](/flowcharts/cluster-garch/) — *Holy Grail*, rigor 8.5 - [The Three-Dimensional Decomposition of Volatility Memory](/flowcharts/the-three-dimensional-decomposition-of-volatility-memory/) — *Holy Grail*, rigor 8.0 - …and 105 more in the [garch family & extensions area](/topics/volatility/) ## Realized volatility & HAR High-frequency measures and the stubborn baseline. - [Multifractality in Bitcoin Realised Volatility: Implications for Rough Volatility Modelling](/flowcharts/multifractality-in-bitcoin-realised-volatility-implications/) — *Holy Grail*, rigor 9.0 - [Geometric Deep Learning for Realized Covariance Matrix Forecasting](/flowcharts/geometric-deep-learning-for-realized-covariance-matrix-forec/) — *Holy Grail*, rigor 8.0 - [Asymptotic Expansions for High-Frequency Option Data](/flowcharts/asymptotic-expansions-for-high-frequency-option-data/) — *Holy Grail*, rigor 8.0 - [Jump detection in financial asset prices that exhibit U-shape volatility](/flowcharts/jump-detection-in-financial-asset-prices-that-exhibit-u-shap/) — *Holy Grail*, rigor 8.5 - [Efficient Sampling for Realized Variance Estimation in Time-Changed Diffusion Models](/flowcharts/efficient-sampling-for-realized-variance-estimation-in-time-changed-diffusion/) — *Holy Grail*, rigor 8.0 - [High-Frequency Volatility Estimation with Fast Multiple Change Points Detection](/flowcharts/high-frequency-volatility-estimation-with-fast-multiple-change-points-detection/) — *Holy Grail*, rigor 8.0 - …and 61 more in the [realized volatility & har area](/topics/volatility/) ## Implied & options-based What option surfaces know that history doesn't. - [Proof-Carrying No-Arbitrage Surfaces: Constructive PCA-Smolyak Meets Chain-Consistent Diffusion with c-EMOT Certificates](/flowcharts/proof-carrying-no-arbitrage-surfaces-constructive-pca-smoly/) — *Holy Grail*, rigor 8.5 - [A Risk-Neutral Neural Operator for Arbitrage-Free SPX-VIX Term Structures](/flowcharts/a-risk-neutral-neural-operator-for-arbitrage-free-spx-vix-te/) — *Holy Grail*, rigor 8.5 - [Fast reliable pricing and calibration of the rough Heston model](/flowcharts/fast-reliable-pricing-and-calibration-of-the-rough-heston-mo/) — *Holy Grail*, rigor 8.5 - [Heath-Jarrow-Morton meet lifted Heston in energy markets for joint historical and implied calibration](/flowcharts/heath-jarrow-morton-meet-lifted-heston-in-energy-markets-for/) — *Holy Grail*, rigor 8.0 - [Machine Learning Methods for Pricing Financial Derivatives](/flowcharts/machine-learning-methods-for-pricing-financial-derivatives/) — *Holy Grail*, rigor 8.0 - [Risk-Sensitive Option Market Making with Arbitrage-Free eSSVI Surfaces: A Constrained RL and Stochastic Control Bridge](/flowcharts/risk-sensitive-option-market-making-with-arbitrage-free-essv/) — *Holy Grail*, rigor 7.8 - …and 242 more in the [implied & options-based area](/topics/volatility/) ## Rough & stochastic volatility The continuous-time frontier. - [Dynamic Skewness in Stochastic Volatility Models: A Penalized Prior Approach](/flowcharts/dynamic-skewness-in-stochastic-volatility-models-a-penalize/) — *Holy Grail*, rigor 8.5 - [A multi-factor model for improved commodity pricing: Calibration and an application to the oil market](/flowcharts/a-multi-factor-model-for-improved-commodity-pricing-calibra/) — *Holy Grail*, rigor 8.5 - [Stochastic Volatility Modelling with LSTM Networks: A Hybrid Approach for S&P 500 Index Volatility Forecasting](/flowcharts/stochastic-volatility-modelling-with-lstm-networks-a-hybrid/) — *Holy Grail*, rigor 8.5 - [A unified theory of order flow, market impact, and volatility](/flowcharts/a-unified-theory-of-order-flow-market-impact-and-volatility/) — *Holy Grail*, rigor 8.0 - [End-to-End Large Portfolio Optimization for Variance Minimization with Neural Networks through Covariance Cleaning](/flowcharts/end-to-end-large-portfolio-optimization-for-variance-minimiz/) — *Holy Grail*, rigor 8.0 - [From rough to multifractal multidimensional volatility: A multidimensional Log S-fBM model](/flowcharts/from-rough-to-multifractal-multidimensional-volatility-a-mu/) — *Holy Grail*, rigor 7.5 - …and 126 more in the [rough & stochastic volatility area](/topics/volatility/) ## ML volatility forecasting Networks and hybrids against the classical baselines. - [Neural Lévy SDE for State--Dependent Risk and Density Forecasting](/flowcharts/neural-l-vy-sde-for-state-dependent-risk-and-density-forecasting/) — *Holy Grail*, rigor 8.5 - [Constructing Time-Series Momentum Portfolios with Deep Multi-Task Learning](/flowcharts/constructing-time-series-momentum-portfolios-with-deep-multi/) — *Holy Grail*, rigor 8.5 - [Deep Learning vs. Statistical Models for Multi-Horizon Price Forecasting of Second-Hand Electronics: A Systematic Benchmark](/flowcharts/deep-learning-vs-statistical-models-for-multi-horizon-price-forecasting-of/) — *Holy Grail*, rigor 9.0 - [A Framework for Predictive Directional Trading Based on Volatility and Causal Inference](/flowcharts/a-framework-for-predictive-directional-trading-based-on-vola/) — *Holy Grail*, rigor 7.0 - [Partial multivariate transformer as a tool for cryptocurrencies time series prediction](/flowcharts/partial-multivariate-transformer-as-a-tool-for-cryptocurrenc/) — *Holy Grail*, rigor 8.0 - [Adaptive Nesterov Accelerated Distributional Deep Hedging for Efficient Volatility Risk Management](/flowcharts/adaptive-nesterov-accelerated-distributional-deep-hedging-fo/) — *Holy Grail*, rigor 7.0 - …and 35 more in the [ml volatility forecasting area](/topics/volatility/) --- **Evaluation lens for this field:** see the [guides index](/guides/) · **Full hub:** [/topics/volatility/](/topics/volatility/) ---------------------------------------------------------------------- # Alpha, Beta, and Alternative Risk Premia: The Difference That Prices Everything URL: https://thequant.space/guides/alpha-beta-alternative-risk-premia/ Section: Guides Date: 2026-09-07 Description: The alpha/beta/ARP taxonomy: what each actually is, the regression that sorts any return stream, why the boundaries move over time, and what each category should cost. **Not an implementation claim.** This is a classification framework — the one that determines what any return stream (a paper's, a fund's, your own) is *worth*, since the three categories command fee structures an order of magnitude apart. **Definitions.** **Beta**: return from passive market exposure — buyable for basis points. **Alternative risk premia (ARP)**: returns from systematic, documented style exposures — [momentum](/guides/cross-sectional-vs-time-series-momentum/), value, carry, low-vol — implementable by rules, priced between index funds and hedge funds, and *arguably compensation for bearing identifiable risks* ([crash risk for momentum, drawdown regimes for carry](/guides/regime-dependence/)). **Alpha**: return unexplained by market and documented styles — scarce, decaying, and the only thing meriting performance fees. The practical point: these are *residual categories from a regression*, not marketing labels, and the regression is runnable by anyone. ## The sorting regression ```text r_strategy − r_f = α + β_mkt·MKT + Σ β_i·ARP_i + ε ``` Regress the stream on market plus a standard style set (the [factor-hub canon](/topics/factor-investing/): value, momentum, quality, carry, low-vol as relevant to asset class). Read it in order: **R² and the betas** tell you what the stream *is* — a "market-neutral" fund loading 0.6 on momentum is an ARP product; **α's magnitude and [its standard error](/guides/confidence-intervals-strategy-research/)** tell you whether anything merits the name (α with a t of 1.3 on four years is [a shrug, not a discovery](/guides/statistical-vs-economic-significance/)); and **ε's behavior in stress months** tells you whether the "alpha" is actually [an unlisted tail-risk premium](/guides/high-sharpe-ratio-not-investable/#3-volatility-isnt-risk--skew-is) — smooth ε that gaps down in crises is short-vol exposure the factor set didn't span. The unglamorous corollary: run this on *your own strategies*. [Stat-arb books](/guides/statistical-arbitrage-signal-to-portfolio/#step-1-neutralization--deciding-what-youre-betting-on) that never met a factor regression are usually momentum-and-short-vol in a trench coat. ## The migration: alpha decays into beta The boundary moves one direction. Yesterday's alpha (size effects in the '80s, momentum in the '90s, [imbalance signals](/guides/limit-order-book-imbalance/) in the 2000s) becomes today's documented premium becomes tomorrow's cheap product — the [post-publication decay literature](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer) is this migration measured. Two consequences: **any alpha claim carries an implicit expiry**, and the diligence question is what protects the residual (capacity limits, infrastructure, data access, [execution quality](/guides/transaction-costs-slippage-market-impact/)) rather than whether it exists today; and **the factor set is era-relative** — α against the 1995 factor set contains things that are ARP against 2026's. Every alpha estimate is conditional on a benchmark vintage, which papers rarely state and readers should. ## What each category should cost — the fee arbitrage lens The taxonomy's cash value is fee diligence: beta at active fees is the classic mis-sale; ARP at 2-and-20 is the hedge-fund industry's documented embarrassment (large fractions of aggregate hedge-fund returns regress onto mechanical style portfolios); and true α commands whatever the market bears, *after* the regression proves it. The same lens prices research: a [paper](/guides/evaluate-trading-backtest/) whose strategy returns regress heavily onto known premia has documented an implementation, not a discovery — worth reading, not worth its claimed novelty. ## Common ways the taxonomy fails its users in production - **Benchmark gaming**: α manufactured by regressing against an incomplete factor set; add the obvious missing premium and it vanishes. (Reader's move: add it.) - **Nonlinear exposures slip the net**: option-selling and [liquidity provision](/guides/market-making-mechanics/) show low *linear* betas and severe conditional exposure — the regression needs stress-month inspection, not just full-sample fits. - **ARP implementation slippage**: the documented premium and your investable version differ by [the tradability gates](/guides/what-makes-a-factor-tradable/); budgeting the paper premium as "beta you can buy" double-counts what implementation eats. - **Alpha-at-scale delusion**: α measured at paper size, charged at fund size — [capacity](/guides/capacity-constraints/) converts α to β through impact. ## Questions to ask of any return stream 1. What's left after the standard regression — and what's the *stress-conditional* residual? 2. Which factor-set vintage was the benchmark, and would today's set absorb more? 3. What structurally protects the residual from migration? 4. Is the fee proportional to the α term, or to the whole stream? ## What our scoring captures Papers that run this decomposition on their own strategies — reporting factor-adjusted alphas with honest benchmarks — occupy the [rigor axis's](/score-guide/) upper reaches; "we beat the market" without a factor regression is the genre's most reliable low-rigor tell. The taxonomy is also the [quadrant system's](/score-guide/) economic shadow: Street Traders papers mostly document ARP; genuine α claims concentrate — and deserve maximum scrutiny — among the Holy Grail's top scorers. --- **The premia themselves:** [factor tradability →](/guides/what-makes-a-factor-tradable/) · [momentum's families →](/guides/cross-sectional-vs-time-series-momentum/) · **The residual's price:** [capacity →](/guides/capacity-constraints/) ---------------------------------------------------------------------- # Auditing a Monte Carlo: Method Notes from a Leveraged-ETF Simulation URL: https://thequant.space/guides/auditing-a-monte-carlo/ Section: Guides Date: 2026-09-06 Description: Every technique used to take four Monte Carlo simulation scripts apart, find what they were actually measuring, and rebuild them — with the real numbers each step produced. Four Monte Carlo scripts, one leveraged ETF pair (LUNR and its 2× daily fund LUNL), 879 + 149 daily bars of data — and a headline probability that moved from **46.8% to 30.9%** once the scripts were audited. These are the method notes: every technique used to take the simulations apart, find what they were actually measuring, and rebuild them, with the real numbers each step produced. The subject is one ticker pair; the techniques apply to any simulation you're about to trust. ## 1. Establish what your data actually is Every downstream number inherits whatever your data source really means. Three things get assumed constantly and are worth five minutes of checking each. ### 1.1 Verify field semantics against a finer dataset I needed to know whether the daily bar's `open` was the 9:30 regular-session open or the first pre-market print — the entire overnight-gap analysis hinges on it. Rather than read documentation, I pulled minute bars for the same day and compared: ```text 2026-08-13: daily-agg open 14.2900 | first pre-mkt 17.1500 | 09:30 RTH open 14.2900 daily-agg close 17.5600 | 15:59 RTH close 17.5600 | last tick 17.6799 ``` Daily open equals the 09:30 bar exactly, and the close is the 16:00 close, not the 19:57 extended-hours tick. **Technique: cross-validate a coarse dataset against a finer one you can reconcile.** One request settles a question that guesswork would leave open all project. (This is the same discipline our [market-data guide](/guides/market-data-vendors/#3-what-do-the-timestamps-mean) demands of timestamps.) ### 1.2 Know precisely what "adjusted" adjusts The API's `adjusted=true` handles splits only — *not* dividends. For a distributing fund, price-only returns then understate total return, and that gap would masquerade as a cost. Since I was about to claim a large unexplained drag, I had to rule this out, so I queried the dividends endpoint directly: none for either ticker. Only then was the drag attributable to costs rather than measurement. ### 1.3 Your sample period is a modelling assumption The same ticker gives wildly different parameters depending on the window: | LUNR window | ann. vol | kurtosis | worst day | |---|---:|---:|---:| | full history (879 bars, from 2023-02) | 1.962 | 203.86 | −75.2% | | trailing 252 days | 1.107 | 2.73 | −15.9% | Full history spans the 2023 SPAC period with a +251% single day. Neither window is "right" — but choosing one silently chooses your tail behaviour, and that choice deserves to be an explicit, named parameter rather than a hardcoded start date.
Defect found One script requested history from `2019-01-01` for a stock that first traded in February 2023, and fitted parameters on the resulting SPAC-contaminated sample.
## 2. Measure the relationship, don't assume it The scripts assumed the ETF returned exactly 2× the stock minus a stated expense ratio. That is a hypothesis, and there were 148 days of data sitting right there to test it against. ### 2.1 What an OLS fit actually buys you *y*ₜ = β·*x*ₜ + α + εₜ Four separate deliverables from one line of code. **β** is realised exposure. **α** is systematic drift the exposure doesn't explain — costs, in this case. **sd(ε)** is tracking error. **R²** tells you whether the linear story is the whole story. ```python beta, alpha = np.polyfit(x, y, 1) resid = y - (beta * x + alpha) r2 = np.corrcoef(x, y)[0, 1] ** 2 # beta 1.9846 alpha -0.001735/day resid_sd 0.007105/day R2 0.9978 ``` β is essentially 2 and R² is 0.998, so the fund tracks superbly. But α annualises to **−43.7%/yr** against the 1.31%/yr the script assumed — a factor of 33. ### 2.2 Annualisation: ×252 for means, ×√252 for volatility Means add over time, so a daily mean scales linearly. Variance adds over independent periods, so *standard deviation* scales with the square root. Mixing these up is the most common unit error in finance code. ```python mean_annual = r.mean() * 252 vol_annual = r.std(ddof=1) * np.sqrt(252) # sqrt, because VARIANCE is what adds ``` Use `ddof=1` for a sample standard deviation — the divisor is *n*−1, since estimating the mean from the same data costs one degree of freedom. ### 2.3 Attack a surprising number before believing it −43.7%/yr was implausible enough that reporting it unchallenged would have been negligent. Four independent attacks, each capable of killing the finding: | Attack | What would kill the finding | Result | |---|---|---| | Month-by-month stability | concentrated in one month | every month, −27% to −67% | | Trim 5 largest moves | driven by outliers | −42.1%, barely moves | | Liquidity check | stale closing prints | $7.2M/day median volume | | Dividend check | distributions not in price | none on record | **Technique: for each explanation that would make your finding boring, run the test that would reveal it.** Surviving four such tests is what converts a number into a claim. ### 2.4 Variance decomposition Splitting close-to-close returns into overnight (4pm→9:30) and intraday (9:30→4pm) legs localises where risk actually lives: for LUNR, **14.2% overnight, 72.6% intraday**. Note the shares don't sum to 100% — because var(*a*+*b*) = var(*a*) + var(*b*) + 2cov(*a*,*b*), and simple returns compound rather than add. Always work in logs when you want additivity, and expect a cross-term when you don't. ### 2.5 The encompassing regression — how to settle an argument Two competing hypotheses about when the fund holds exposure. The weak approach fits each separately and compares R². The strong approach **nests both in one model and lets them compete for the same variance**: LUNLcc = *b*₁·LUNR_overnight + *b*₂·LUNR_intraday + *a* A fund holding nothing overnight must produce *b*₁ = 0. Measured: ```python X = np.column_stack([np.ones(n), overnight, intraday]) coef, *_ = np.linalg.lstsq(X, y, rcond=None) se = np.sqrt(np.diag(np.linalg.inv(X.T @ X) * resid.var(ddof=3))) # b1 (overnight) = 2.0175 +/- 0.0450 b2 (intraday) = 1.9832 +/- 0.0236 ``` The t-statistic against the null is *t* = (estimate − null) / SE = 2.0175/0.0230 = **87.8 standard errors** from zero. That is not a close call, and it is a far more decisive instrument than eyeballing two R² values.
4:00pm close 9:30am open 4:00pm close overnight + pre-market regular session reset at NAV 2× exposure held continuously 9:30–4pm only flat — predicts β = 0 2× exposure measured β = 2.02 β = 1.98 R² = 0.958 R² = 0.986 the overnight window is where the two hypotheses disagree — and the data answers it
Both windows measure β ≈ 2. The dashed box marks the only place the hypotheses differ; a fund flat overnight cannot produce β = 2.02 with R² = 0.958 against a gap it isn't exposed to.
## 3. Learn the arithmetic of the instrument ### 3.1 Work in log returns With *r*_log = ln(1 + *r*_simple), returns become additive over time, which makes drift, variance and horizon scaling all linear. It has a second benefit that matters here: **generating in log space makes a return below −100% impossible by construction**, so you never need a floor to patch up a broken sampler. ```python log_r = mu + sd * z # generate here simple = np.expm1(log_r) # convert out; always > -1 ``` ### 3.2 Volatility drag, derived For geometric Brownian motion, Itô's lemma gives the log dynamics of the stock, and the fund holds *L*× that exposure: > d ln *S* = (μ − ½σ²) d*t* + σ d*W* > d ln Λ = (*L*μ − ½*L*²σ²) d*t* + *L*σ d*W* Subtract *L*× the first from the second and the μ and dW terms cancel, leaving pure drag: > **drag = −½σ²(*L*² − *L*) = −σ² for *L* = 2** No μ term: **drag is unrelated to direction, and quadratic in leverage.** A 3× fund suffers three times the drag of a 2× fund, not 1.5×. At LUNR's σ = 1.107 that is −1.225/yr at 2× and −3.676/yr at 3×. ### 3.3 The subtlety: it's mean-square, not variance The discrete expansion is −½·E[*r*²]·(*L*²−*L*), and E[*r*²] = σ² + μ². Those coincide only when μ ≈ 0, which is why the textbook writes σ². Ten consecutive −5% days have *zero variance*, yet: ```text ln(2x fund) = -1.05361 2 * ln(underlying) = -1.02587 gap -0.02774 predicted: -0.5 * mean(r**2) * (L**2 - L) * 10 = -0.02500 ``` The fund lags 2× the log return even with constant returns. Drag is driven by the mean square of the daily move, not by variance in the statistical sense. ### 3.4 Two benchmarks — the source of nearly all confusion "Does a leveraged ETF beat 2×?" has no answer until you say *which* 2×: | 10 days | fund | L× log return | L× simple return | |---|---:|---:|---:| | trend: −5% ×10 | 0.3487 | 0.3585 | **0.1975** | | chop: ±5% ×10 | 0.9510 | 0.9753 | 0.9751 | Against **L× the log return** the fund always lags — drag is unavoidable. Against **L× the simple return** the fund *wins* in trends and loses in chop, because sustained moves de-lever you into losses. Both statements are true; they are answers to different questions. ### 3.5 Terminal value is a product, so ordering cannot matter ∏(1 + *L*·*r*ᵢ) is invariant under permutation. Multiplication commutes. Testing it on the fund's real 148 returns: ```python terms = [np.prod(1 + rng.permutation(y)) for _ in range(20000)] # true order 0.4592145015x min 0.4592145015 max 0.4592145015 spread 2.05e-15 ``` The −½σ²(*L*²−*L*) formula contains only a variance term — no autocorrelation, no sequencing. That is the same fact stated analytically. ### 3.6 ...but the path is entirely about ordering
$10 break-even $3 stop 0 5 10 trading day peak $12.24 trough $2.49 both end $4.011 B: +5%×5 then −10%×5 A: −10%×5 then +5%×5
The same ten daily returns, reordered, both starting at $7.60. Identical endpoint to fifteen decimal places — but B trades above the break-even for three days while A never exceeds its start and falls through a $3 stop. Max drawdown happens to be identical (−67.23%) in both: it is the same contiguous run of five −20% fund days either way.
So ordering drives drawdown timing, barrier touches and stop-outs — everything you would actually trade against — and nothing about where you land if you hold. Report both kinds of statistic, and know which kind each one is. (Our [equity-curve simulator](/tools/equity-curve-simulator/) makes the same point interactively: the fan is the strategy; one path is one draw.) ## 4. Generate futures honestly ### 4.1 The generator menu | Generator | Preserves | Cannot produce | |---|---|---| | Gaussian | mean, variance | fat tails, clustering | | Student-*t* | mean, variance, fat tails | clustering | | iid bootstrap | exact empirical marginal | clustering, unseen moves | | Block bootstrap | marginal + short-run clustering | unseen moves, long memory | | Stationary bootstrap | as above, no fixed seams | unseen moves | The **stationary bootstrap** (Politis–Romano) draws geometric block lengths rather than fixed ones, so no artificial boundary recurs every *b* days. Cheap to implement and strictly better behaved than fixed blocks. ### 4.2 Verify the sampler does what you claim Don't assert that block sampling preserves volatility clustering — measure the autocorrelation of |*r*|: ```text |r| autocorrelation: iid -0.001 block 0.394 ``` ### 4.3 Clustering acts through dispersion, not sequence Given §3.5, how can clustering matter at all if order doesn't? Because it changes the *multiset* each path draws. Block sampling lets one path pull an entire calm stretch and another an entire violent one, widening the spread of per-path variance — and drag keys off variance: ```text per-path realised vol LUNL median LUNL p05 iid bootstrap mean 1.094 sd 0.149 7.744 1.399 block bootstrap mean 1.085 sd 0.201 8.290 1.040 ``` Same marginal distribution, materially fatter loss tail. That is the real argument for block over iid — and it is a dispersion argument, not an ordering one. ### 4.4 A bootstrap cannot invent a tail it never saw ```text worst day in the 252d sample -15.93% worst day the block bootstrap drew -15.93% # identical, by construction worst day Student-t (df=3) drew -99.99% P(any day <= -50%, a 2x wipeout): bootstrap 0.0000 student-t 0.0112 ``` **Every resampling method reports exactly zero wipeout probability — a property of your sample, not of the world.** For tail questions you need a parametric generator that can extrapolate. Judgement still required in the other direction: Student-*t* with df = 3 produced a −99.99% day, which is its own kind of nonsense. Neither is truth; know which failure mode you've chosen. ### 4.5 Drift calibration: mean and median are different targets For GBM, E[*S*_T] = *S*₀·e^(μT) but median = *S*₀·e^((μ−½σ²)T). At σ = 1.107 over 60 days the median sits **13.6% below** the mean. Setting μ = ln(target/*S*₀)/T therefore aims the *mean* at your target while the median — the number people actually read — lands well short. The robust fix avoids the algebra entirely: simulate first, then shift every log return by a constant so the finished sample lands where you want. Exact for either statistic, and immune to whatever else is in your generator. ```python log_r = np.log1p(sim_rets) shift = (np.log(target / S0) - np.median(log_r.sum(axis=1))) / n_days tilted = np.expm1(log_r + shift) # median now lands exactly on target ```
Defect found Adding jumps with a non-zero mean adds unintended drift, because E[e^J] ≠ 1. Unless you subtract the compensator λ(E[e^J]−1), the calibration silently breaks: one script's bull scenario carried **+25.3%/yr** of drift nobody asked for.
### 4.6 Absorbing barriers Flooring a daily return at −0.999 leaves a fund that lost 99.9% alive and able to rally back, inflating the right tail with paths that in reality would have been liquidated. Once dead, keep it dead: ```python dead = np.cumsum(r <= -1.0, axis=1) > 0 nav[:, 1:] = np.where(dead, 0.0, nav[:, 1:]) ``` ## 5. Quantify your own uncertainty A Monte Carlo result is a sample statistic with its own sampling error. Printing "43.33%" without an error bar hides whether the third digit means anything. ### 5.1 Standard error of a simulated probability SE = √( *p*(1−*p*) / *n* ) Verified empirically — 60 independent runs at *n* = 5,000: ```text observed sd of p_hat across runs 0.00674 formula sqrt(p(1-p)/n) 0.00706 # matches ``` The √*n* means **halving your error costs 4× the paths**. At 200,000 paths a probability near 50% carries ±0.22pp at two standard errors — which tells you a 3pp difference between scenarios is real and a 0.1pp difference is noise. ### 5.2 Quantiles have no closed form — bootstrap them ```python idx = rng.integers(0, len(x), size=(200, len(x))) se = np.std(np.quantile(x[idx], q, axis=1), ddof=1) ``` Resample your own output, recompute the statistic, take the spread. The same trick works for any statistic you can compute but not analyse. ### 5.3 Variance reduction: antithetic variates For every draw *z*, also use −*z*. Symmetric error cancels in the pair rather than averaging down slowly, and it costs nothing: ```python half = (n + 1) // 2 z = rng.standard_normal((half, d)) return np.vstack([z, -z])[:n] # sd of mean terminal, plain 0.01109 # sd of mean terminal, antithetic 0.00603 # 1.8x tighter, same cost ``` Its companion is **common random numbers**: when comparing scenarios, drive both with the same seed so the difference you measure is the assumption, not the sampling. ### 5.4 Report the sensitivity, not the point estimate The most useful output of the whole exercise wasn't a probability — it was finding which assumption owned it: | Underlying drift | Cost model | P(break-even) | |---|---|---:| | historical (+141%/yr) | prospectus 1.31%/yr | **46.8%** | | historical | fitted −43.7%/yr | 43.6% | | zero | prospectus | 34.1% | | zero | fitted | **30.9%** | A 16-point swing from two assumptions the original scripts never surfaced. The bootstrap silently inherited the stock's trailing +141%/yr drift — a strongly bullish input nobody had chosen on purpose. **Any parameter that can move your answer that far belongs in a table, not buried in a default.** ## 6. Prove the code does what you claim ### 6.1 Test invariants, not outputs Asserting that a simulation returns 30.9% just freezes today's bug. Assert *properties* that must hold for any correct implementation: ```text ok view tilt lands the requested statistic exactly ok wipeout absorbing (dead path stays at 0 through later rallies) ok backtest aligned (exact 2x series reproduces to rmse=1.73e-16) ok all 5 generators match historical log-vol and stay above -100% ok |r| autocorr: iid=-0.001 vs block=0.394 (clustering preserved) ok antithetic normals cancel mean error exactly (sd 2.1e-19 vs 7.1e-04) ``` The third is the strongest pattern available: **feed the model data it should reproduce perfectly** — a synthetic exact-2× series — and demand machine precision. Any alignment or convention error shows up instantly. ### 6.2 Cross-check simulation against closed form Wherever theory gives an answer, make the simulator reproduce it. Comparing measured drag against −½σ²(*L*²−*L*): ```text L=2.0: theory -1.2252/yr simulated -1.1882/yr L=3.0: theory -3.6757/yr simulated -3.6774/yr ``` This validates the compounding, the sign conventions and the leverage handling in one shot. ### 6.3 Off-by-one errors hide in cumulative products
Defect found `cumprod(1+r)[0]` already contains day one's return, so comparing it against a NAV normalised to 1.0 on that same day offsets the two series by a day. Both original scripts did this, inflating reported RMSE from **0.0136 to 0.2065** — a 15× error in the one number meant to prove the model worked.
```python model_nav = np.r_[1.0, np.cumprod(1.0 + model_daily[1:])] # both start at 1.0 ``` Always verify that two series you compare start on the same date *at the same value*. ### 6.4 Vectorise — speed buys statistical quality ```text scalar Python double loop ~28 s (100,000 paths x 126 days) vectorised numpy ~0.2 s # 140x ``` This isn't cosmetic. One script ran 100 paths because more was too slow, and 100 paths cannot support a 5th percentile. Fast code is what lets you afford 200,000 paths, a sensitivity table and honest error bars. ### 6.5 Make the run reproducible Seed the generator explicitly, cache fetched data to disk, keep configuration in one place, take secrets from the environment, and emit assumptions alongside results so the output can be re-derived months later. ## The through-line Nearly every finding came from the same reflex: asking what a number was actually measuring rather than what it was labelled. Four questions did most of the work. 1. **Is this parameter assumed or measured?** The 1.31%/yr cost was assumed; measuring it gave −43.7%. 2. **What sample produced this, and what does that sample exclude?** The window choice set the volatility, the tails, and whether wipeout risk existed at all. 3. **Which statistic is this — mean, median, or something else?** Drift calibration, benchmark comparison and drag all turn on this distinction. 4. **How would I know if this were wrong?** If no test could fail, it isn't validation. The scripts weren't badly written. They were confidently answering a question slightly different from the one being asked — which is the normal failure mode of quantitative work, and why the audit is worth as much as the model. It's the same standard we hold [papers to](/score-guide/), applied to our own code — and the reason the [production checklist](/guides/production-checklist/) treats every unverified number as a lie until an independent mechanism confirms it. ---------------------------------------------------------------------- # Capacity Constraints in Quantitative Strategies URL: https://thequant.space/guides/capacity-constraints/ Section: Guides Date: 2026-09-07 Description: How to estimate a strategy's capacity: the impact arithmetic, the capacity hierarchy by strategy type, self-competition effects, and why capacity is the number papers never report. **Not an implementation claim.** Capacity analysis is the arithmetic that converts "does it work?" into "does it matter?" — the question that [separates academic from implementable](/guides/academic-vs-implementable/#the-five-implementability-tests) and the number conspicuously absent from most published backtests. **Definition.** A strategy's **capacity** is the asset level at which expected net returns fall to some threshold (zero, or an opportunity-cost hurdle) because the strategy's own trading moves prices against it. It exists for every strategy, because [impact grows with participation](/guides/transaction-costs-slippage-market-impact/#the-four-components-separated) while edges don't. ## The estimation arithmetic A serviceable first-order capacity model needs four numbers you can estimate from any honest backtest: ```text per-rebalance participation = (AUM × turnover_per_rebalance) / (universe ADV captured) impact per rebalance ≈ k · σ · √participation (square-root law) capacity: solve for AUM where annualized impact ≈ gross edge ``` Worked example: a daily-rebalanced strategy trading 50 liquid names ($50M combined relevant ADV), 10% daily turnover, gross edge 6%/yr, daily σ ≈ 1.5%, k ≈ 0.5. At $100M AUM: participation = 20% of ADV — deep in impact territory; the square-root law prices each rebalance at ~34 bps, ~85%/yr annualized — absurd, the strategy is *far* over capacity. Iterate downward and the viable AUM lands near single-digit millions. The point isn't the constants (calibrate `k` to your own fills); it's that **five minutes of arithmetic bounds what a paper's Sharpe is worth** — and routinely reveals published strategies to be [salary-sized, not fund-sized](/guides/high-sharpe-ratio-not-investable/#2-it-doesnt-scale). ## The capacity hierarchy Capacity is set by holding period × universe liquidity — which orders the strategy world: | Strategy type | Typical capacity order | Why | |---|---|---| | [HFT / market making](/guides/market-making-mechanics/) | $1–100M | Edge per trade ~ticks; participation caps are brutal | | [Tick/LOB signals](/guides/limit-order-book-imbalance/) | similar | Fast decay forces urgency; urgency is impact | | [Stat-arb / short-horizon equity](/guides/statistical-arbitrage-signal-to-portfolio/) | $100M–low $B | Breadth helps; turnover hurts | | [Momentum-family factors](/guides/cross-sectional-vs-time-series-momentum/) | $B-scale, contested | Turnover mid; crowding severe | | Low-turnover value/quality tilts | $10B+ | Slow trading in liquid names | | Beta | ~unbounded | Nothing to compete away | The diagonal is the industry's structure: high-Sharpe/low-capacity strategies fund proprietary desks; low-Sharpe/high-capacity ones become products. A paper's implied position on this diagonal tells you who could possibly care about its result — [and at what fee](/guides/alpha-beta-alternative-risk-premia/#what-each-category-should-cost--the-fee-arbitrage-lens). ## Capacity is shared: the crowding correction Your capacity model prices *your* participation; the market bills *aggregate* participation in the trade. Published strategies are shared strategies, so effective capacity is divided among everyone running the signal — the mechanism behind [post-publication decay](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer) and the reason [crowding proxies](/guides/what-makes-a-factor-tradable/#gate-4-crowding--the-premiums-reflexive-component) belong in capacity estimates. Corollary: capacity estimated from a backtest (you alone) is an upper bound that publication itself erodes. ## Common ways capacity fails in production - **The backtest's fills assumed away the question**: unconstrained position sizes in [thin names](/guides/evaluate-factor-investing-paper/#the-factor-paper-checklist) mean the reported Sharpe was earned at an AUM near zero. - **Scaling changed the strategy**: liquidity caps and participation limits added at size reshape the portfolio until realized returns decorrelate from the research version — you're running a different, undertested strategy. - **Capacity consumed by success**: the strategy works, AUM grows, returns decay on schedule, and the decay gets [misdiagnosed as regime or alpha loss](/guides/regime-dependence/#the-live-trading-diagnosis-problem) instead of self-impact. Track return-vs-AUM explicitly. - **Exit capacity ≠ entry capacity**: positions accumulated patiently must sometimes exit urgently; [stress liquidity](/guides/high-sharpe-ratio-not-investable/#6-costs-and-slippage-are-convex-in-urgency) is the binding constraint, and it's a fraction of calm-market ADV. ## Questions to ask a paper 1. What AUM do the reported positions imply against the universe's ADV? 2. Does performance survive a 5%-of-ADV participation cap? 3. Where does the strategy sit on the hierarchy, and does its Sharpe make sense for that seat? 4. If this were published and adopted, what's the *shared* capacity? ## What our scoring captures Capacity reporting is rare enough to be a top-decile [rigor](/score-guide/) signal — papers with participation caps, liquidity-bucketed results, or explicit capacity estimates announce that the authors have met real order books. The hierarchy also explains a scoring pattern: [HFT-hub](/topics/high-frequency-trading/) papers can be simultaneously high-rigor and low-relevance to most readers — sound science about seats you can't sit in. The score certifies the evidence; the arithmetic above tells you if the seat is yours. --- **The impact law:** [transaction costs →](/guides/transaction-costs-slippage-market-impact/) · **The consequence:** [why high Sharpe ≠ investable →](/guides/high-sharpe-ratio-not-investable/) · [factor tradability →](/guides/what-makes-a-factor-tradable/) ---------------------------------------------------------------------- # Combinatorially Symmetric Cross-Validation (CSCV) Explained URL: https://thequant.space/guides/cscv-explained/ Section: Guides Date: 2026-09-07 Description: CSCV and the Probability of Backtest Overfitting (PBO) explained: the block-combination construction, pseudocode, how to read PBO, and the method's honest limitations. **Definition.** Combinatorially Symmetric Cross-Validation (CSCV), from Bailey, Borwein, López de Prado and Zhu's work on backtest overfitting, estimates the **Probability of Backtest Overfitting (PBO)**: the probability that the strategy configuration you selected *because* it won in-sample would rank in the bottom half out-of-sample. Where the [deflated Sharpe ratio](/guides/deflated-sharpe-ratio/) benchmarks your best result against luck's expected maximum, CSCV directly *measures* how unstable your selection is — no distributional assumptions about returns required, only your own trial matrix. ## The construction Start from what a search process actually produces: a matrix **M** of returns with T rows (time) and N columns (every configuration you evaluated — the [honest N](/guides/statistical-vs-economic-significance/#why-the-usual-significance-bar-is-broken-in-finance), not the flattering one). 1. **Partition** the T rows into S equal, contiguous blocks (S = 16 is a common choice), preserving temporal order within blocks. 2. **Form all combinations** of S/2 blocks as the in-sample set; the complementary S/2 blocks are out-of-sample. That's C(16, 8) = 12,870 symmetric splits — every split has a mirror, so in- and out-samples are statistically interchangeable, unlike a single chronological split. 3. **For each split**: rank all N configurations in-sample; take the in-sample winner n*; find its *out-of-sample rank* among all N; convert to a relative rank ω ∈ (0,1). 4. **PBO** = the fraction of splits where ω < 0.5 — the in-sample winner landed in the OOS bottom half. The logit λ = ln(ω/(1−ω)) across splits gives a full distribution, not just the headline probability. ```python from itertools import combinations import numpy as np def pbo(M, S=16): # M: T x N matrix of per-period returns blocks = np.array_split(np.arange(len(M)), S) below = 0; splits = list(combinations(range(S), S // 2)) for c in splits: ins = np.concatenate([blocks[i] for i in c]) outs = np.concatenate([blocks[i] for i in range(S) if i not in c]) sr = lambda X: X.mean(0) / (X.std(0) + 1e-12) winner = sr(M[ins]).argmax() w = (sr(M[outs]) < sr(M[outs])[winner]).mean() # OOS relative rank of winner below += (w < 0.5) return below / len(splits) ``` Run it on the [thousand-coin-flips demo](/guides/backtest-overfitting/#the-ten-line-demonstration) and PBO comes out near 0.5 — the in-sample "winner" is a coin toss out-of-sample, exactly what overfitting means. A selection process with real signal pushes PBO toward 0. ## Reading the number - **PBO ≲ 0.2**: selection is finding something persistent; proceed to the [rest of the gauntlet](/guides/evaluate-trading-backtest/). - **PBO ≈ 0.5**: your selection is noise — the best backtest carries no information about OOS ranking. - **PBO > 0.5**: worse than noise; the search is anti-selecting (usually a tell of strong mean-reverting noise fit). Report it alongside the winner's Sharpe: "SR 1.8, PBO 0.11, N = 240 configurations" is a claim; "SR 1.8" is a mood. ## Honest limitations - **It needs the trial matrix.** CSCV audits the configurations you *kept records for* — untracked exploratory trials escape it, which is the same reason the [experiment log](/guides/quant-research-pipeline/#1-idea-intake--the-experiment-log) is the foundational defense. - **Block boundaries leak slightly** with overlapping features/labels; apply the same [purge-and-embargo](/guides/walk-forward-out-of-sample-testing/#why-the-ml-playbook-fails-on-market-data) hygiene at block edges. - **Combinations aren't causal time**: OOS blocks can predate IS blocks, so CSCV measures selection stability, not deployment realism — it complements [walk-forward](/guides/walk-forward-out-of-sample-testing/), never replaces it. - **It says nothing about data biases**: a universe with [survivorship](/guides/survivorship-bias/) or a pipeline with [leakage](/guides/look-ahead-bias-point-in-time-data/) passes CSCV proudly and is still fiction. ## Questions to ask when a paper reports (or should report) PBO 1. Does N include the full search grid, or only the finalists? 2. Is the performance metric in step 3 the same one the paper's conclusions use? 3. With S blocks, is each block long enough to contain a regime, or is the split shredding [regime structure](/guides/regime-dependence/) into confetti? 4. If the paper reports no overfitting diagnostics at all, what does its search-space size imply the PBO plausibly is? ## What our scoring captures — and misses Papers that report PBO, DSR, or disclosed trial matrices earn top marks on the [rigor axis](/score-guide/) — they're also *rare*, which itself is informative about the field. Our score can't compute PBO for a paper (that requires the trial matrix only authors possess), so the score rewards the disclosure, and the [reproduction process](/guides/reproduce-quant-research/) is where you generate the matrix yourself for results that matter. *Reference: Bailey, Borwein, López de Prado & Zhu, "The Probability of Backtest Overfitting" (Journal of Computational Finance).* --- **The family:** [Backtest overfitting →](/guides/backtest-overfitting/) · [Deflated Sharpe →](/guides/deflated-sharpe-ratio/) · [Walk-forward →](/guides/walk-forward-out-of-sample-testing/) ---------------------------------------------------------------------- # Corporate Actions and Adjusted Price Data URL: https://thequant.space/guides/corporate-actions-adjusted-prices/ Section: Guides Date: 2026-09-07 Description: Corporate actions in quant research: how adjustments work, the dividend/total-return distinction, the actions that break backtests (spinoffs, mergers, delistings), and the store-unadjusted principle. Every price series you've ever backtested embeds an **adjustment policy** — decisions about splits, dividends, spinoffs, and mergers — usually made by a vendor, usually undocumented, and usually discovered during a [reconciliation that won't tie out](/guides/reproduce-quant-research/#stage-4--reconcile-match-statistics-before-results). This guide is the mechanics: what each action does to a naive series, and the architecture that keeps you in control. ## The mechanics of adjustment A 2-for-1 split halves the price overnight; unadjusted, that's a −50% "return." Adjustment multiplies all *earlier* prices by the split factor so the series is return-continuous. Two properties follow that everyone eventually relearns: **adjustment rewrites history** (every new action changes all past adjusted prices — an adjusted series is a *current-vintage* object, [knowledge-timed like any other fact](/guides/point-in-time-data-architecture/#schema-by-data-class--where-the-pattern-bites)), and **adjusted prices are relative, not absolute** (a 1990s adjusted price of $0.43 never traded — price-level logic like "stocks under $5" run on adjusted data is [a filter on the future](/guides/look-ahead-bias-point-in-time-data/)). ## Dividends: the total-return distinction Split-adjustment is uncontroversial; dividends fork the road. **Price-only adjusted** series (most free and many paid feeds — [check which you have](/guides/free-vs-paid-market-data/#the-six-dimensions-that-actually-separate-the-tiers)) show the price drop on ex-date as a loss: a long-only equity backtest understates returns by roughly the dividend yield, *compounded* — 2%/yr for decades is not a rounding error. **Total-return adjustment** reinvests dividends into the factor. Neither is wrong; *unstated* is wrong, and strategy-dependent: dividend-capture and [value-factor](/guides/evaluate-factor-investing-paper/) research need the components separated, short positions *owe* the dividends (a [short-leg cost](/guides/what-makes-a-factor-tradable/#gate-3-the-short-side--half-the-premium-most-of-the-problems) price-adjusted series hide), and cross-vendor [reconciliation](/guides/notebook-to-production/#5-logs-metrics-and-the-alert-that-actually-matters) fails whenever two policies meet. ## The actions that break pipelines - **Spinoffs**: holder receives shares of a new entity; the parent gaps down. Price-only series book a loss; total-return handling needs the spinoff's value — and the new entity needs [an identity](/guides/point-in-time-data-architecture/#schema-by-data-class--where-the-pattern-bites) in your universe *from the distribution date*. The classic silent P&L hole. - **Mergers (cash/stock/mixed)**: positions convert to cash, acquirer shares, or both; naive pipelines see a delisting and drop the [terminal value](/guides/survivorship-bias/#building-a-survivorship-clean-backtest) — turning M&A premia into losses. Stock-for-stock needs share-ratio conversion, not price splicing. - **Delistings**: the final return (often severely negative for cause, positive for buyouts) must reach the P&L; the [delisting-return conventions](/guides/survivorship-bias/) are half of survivorship hygiene. - **Rights issues, special dividends, return-of-capital**: each a different factor computation; vendors disagree, and the disagreements cluster in exactly the [small caps where factors live](/guides/interpret-factor-zoo-paper/#what-survives-everyones-methodology). - **Ticker reuse and identity churn**: actions rename and recycle tickers; joins on ticker [splice unrelated companies](/guides/point-in-time-data-architecture/#schema-by-data-class--where-the-pattern-bites) — entity ids or archaeology, choose one. ## The architecture: store unadjusted + actions, adjust at read The [standing principle](/guides/market-data-vendors/#4-corporate-actions-who-does-the-adjusting): keep **unadjusted prices** and a **corporate-actions table** (action type, ex-date, terms, [knowledge time](/guides/point-in-time-data-architecture/#the-core-pattern-two-timestamps-never-one)), compute adjustment factors at query time, per policy. What it buys: auditability (a factor you computed is checkable; a vendor's baked-in one is faith), policy freedom (price-only, total-return, and component views from one store), vintage correctness (adjusted history as-of any date), and cross-vendor reconciliation (compare *actions*, not entangled outputs). Cost: one well-tested adjustment module — which is exactly the kind of code the [known-answer test pattern](/guides/auditing-a-monte-carlo/#61-test-invariants-not-outputs) was made for: a synthetic security with one of every action type, hand-computed truth, machine-precision assertion. ## Common failure modes in the wild - **The 2%/yr phantom drag**: price-only data mistaken for total-return — [we hit exactly this question auditing a leveraged-ETF study](/guides/auditing-a-monte-carlo/#12-know-precisely-what-adjusted-adjusts), where confirming "no dividends" was a required step before attributing drag to costs. - **Backtest-live divergence on ex-dates**: research on adjusted series, production on raw feeds; every ex-date books a fake signal or fake P&L until [reconciliation](/guides/production-checklist/#gate-5--monitoring-and-reconciliation) catches it. - **Vendor factor revisions**: adjustment factors *change* (corrections, late reports); pinned snapshots plus your own factors-from-actions is the defense. - **Split-driven "volatility"**: unadjusted data with unhandled splits injects ±50% returns into [vol estimates and risk models](/guides/microstructure-noise-realized-volatility/) — the loudest version of the quietest problem. ## Questions to ask (of a vendor or a paper) 1. Price-only or total-return — and where is the policy documented? 2. Are actions available as *data*, or only baked into prices? 3. How were spinoffs, mergers, and delisting proceeds handled? 4. Do price-level filters run on unadjusted values? ## What our scoring captures Adjustment-policy disclosure is a small text signal with outsized [rigor](/score-guide/) information: papers stating their dividend handling and delisting conventions almost always got the bigger things right too. Its absence in a long-horizon equity result is a quiet flag — the [ten-minute read's data section](/guides/how-to-read-quant-finance-papers/#the-10-minute-pass-in-the-right-order) is where to catch it. --- **The identity layer:** [PIT architecture →](/guides/point-in-time-data-architecture/) · **The bias family:** [survivorship →](/guides/survivorship-bias/) · **The vendor question:** [market-data guide →](/guides/market-data-vendors/) ---------------------------------------------------------------------- # Cross-Sectional Momentum vs Time-Series Momentum URL: https://thequant.space/guides/cross-sectional-vs-time-series-momentum/ Section: Guides Date: 2026-09-07 Description: The two momentum families compared: construction, crash profiles, evidence bases, and why they are different strategies that happen to share a name. **Not an implementation claim.** Both momentum families are heavily documented and heavily crowded; this guide teaches their *mechanics* — the construction differences that determine risk, evidence standards, and failure modes — not a recommendation to run either. This is the applied companion to [cross-sectional vs time-series predictability](/guides/cross-sectional-vs-time-series-predictability/): momentum is the flagship example of the distinction, big enough to earn its own mechanics. ## Construction, side by side | | Cross-sectional (relative strength) | Time-series (trend) | |---|---|---| | Signal | Asset's return *rank* among peers | Sign of asset's *own* past return | | Position | Long top decile, short bottom | Long/flat/short each asset independently | | Net exposure | ~zero by construction | Structurally directional; varies with regime | | Natural habitat | Single-asset-class universes (equities) | Multi-asset futures ([managed futures](/topics/factor-investing/)) | | Characteristic crash | Momentum crashes: short-leg rips in sharp reversals (2009's is canonical) | Whipsaw at trend turns; gap reversals (2020's V) | | Skew | Strongly negative | Historically positive-ish (long-vol-like in crises) | The skew line is the deepest difference: cross-sectional momentum *sells* crisis-reversal insurance (it's short the recently-crushed names that bounce hardest), while time-series momentum has historically *provided* crisis offset (it's short whatever was falling) — the "crisis alpha" claim of the managed-futures industry. Same word, opposite tail products. ## The evidence bases differ in kind Cross-sectional momentum's evidence: [large effective samples](/guides/cross-sectional-vs-time-series-predictability/#the-worked-example-same-signal-two-strategies-different-worlds) (breadth × time), robust across countries and [zoo studies](/guides/interpret-factor-zoo-paper/#what-survives-everyones-methodology) — one of the few survivors of harsh corrections — but with well-documented [post-publication decay](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer) and brutal crash episodes. Time-series momentum's evidence: centuries-long backtests across asset classes look imposing, but the *effective* sample is trend episodes, not months — [a few dozen independent bets per asset class](/guides/evaluate-trading-backtest/#4-how-many-independent-bets-is-this-really) — and the strategy's returns concentrate in a handful of big-trend years, making its statistics [regime-mixture averages](/guides/regime-dependence/#the-worked-example-the-average-that-never-happens). Hold TSM papers to episode-counting standards; their t-stats flatter. ## Mechanics that decide live behavior - **Lookback/holding grids are trial factories**: 12-1, 6-1, vol-scaled, skip-months — every published variant is a cell in a searched grid; apply [the plateau standard](/guides/robustness-trading-strategy/#1-parameter-robustness--the-plateau-test) and [DSR counting](/guides/deflated-sharpe-ratio/). - **Vol scaling changes the strategy**: sizing positions inverse to volatility (near-universal in modern TSM) embeds a [volatility-targeting layer](/guides/volatility-targeting/) with its own leverage dynamics — and much of "improved" TSM performance attributes to the vol layer, not the trend signal. - **Turnover asymmetry**: CSM rebalances a whole book monthly (heavy [turnover tax](/guides/turnover-factor-returns/)); TSM trades only at signal flips but concentrates its trading in exactly the volatile turns where [costs spike](/guides/transaction-costs-slippage-market-impact/#6-costs-and-slippage-are-convex-in-urgency). ## Common ways each fails in production **Cross-sectional:** the crash arrives via the *short leg* during broad reversals — junk rallies 40% in a quarter while your longs lag; borrow on crushed names vanishes precisely then; and sector-neutralization (the standard crash mitigant) [taxes the gross](/guides/statistical-arbitrage-signal-to-portfolio/#step-1-neutralization--deciding-what-youre-betting-on) every calm month to soften the rare bad one. **Time-series:** years of whipsaw bleed between trends (the strategy's "cost of carry" is psychological and financial); position flips at turns execute into the [worst liquidity](/guides/transaction-costs-slippage-market-impact/); and the crisis-alpha promise is *conditional* on crises trending rather than gapping — a V-reversal punishes both directions in weeks. **Both:** crowding. Momentum is the most-run signal family on earth; its unwinds are correlated events, and your backtest's [regime mix](/guides/regime-dependence/) contains a limited number of them. ## Questions to ask a momentum paper 1. Which family — and if both appear, are results kept separate? 2. What share of TSM profits comes from the top 5 trend episodes? 3. Is the crash episode *in* the sample, and does the paper report performance through it? 4. How much of the improvement over vanilla momentum is the vol-scaling layer? ## What our scoring captures Momentum papers are a rigor-axis showcase: the genre's best work reports crash episodes, subperiods, and costs unprompted, and ranks accordingly in the [factor hub](/topics/factor-investing/); the genre's worst reports a 40-year gross Sharpe and a lookback grid's best cell. The [two-family distinction](/guides/cross-sectional-vs-time-series-predictability/) itself is what to verify first — papers that pool the families have averaged two different risk products into one flattering number. --- **The general distinction:** [cross-sectional vs time-series predictability →](/guides/cross-sectional-vs-time-series-predictability/) · **The layers on top:** [volatility targeting →](/guides/volatility-targeting/) · [turnover →](/guides/turnover-factor-returns/) ---------------------------------------------------------------------- # Cross-Sectional vs Time-Series Predictability: Two Different Claims URL: https://thequant.space/guides/cross-sectional-vs-time-series-predictability/ Section: Guides Date: 2026-09-07 Description: The difference between cross-sectional and time-series predictability in finance: what each claim means, how evidence standards differ, and why conflating them produces phantom strategies. **Definition.** **Cross-sectional predictability**: a signal ranks assets against each other — high-signal names outperform low-signal names, market direction unknown and hedged away by the long-short construction. **Time-series predictability**: a signal forecasts an asset against its own history — direction itself, no hedge from construction. "Momentum works" means *both* claims to different papers (relative winners keep winning; assets with positive 12-month returns keep rising), and they are so different in mechanics and evidence that treating them as one is a category error with a P&L attached. ## The worked example: same signal, two strategies, different worlds Take 12-month past return on 500 stocks. **Cross-sectional use**: rank monthly, long the top decile, short the bottom. Every month contributes ~500 ranked observations; the market factor nets out; the classic failure is the **momentum crash** — the short leg (junk that cratered) rips upward in sharp reversals, delivering brutal negative skew while beta stays near zero. **Time-series use**: hold each asset long when its own 12-month return is positive, else flat/short. Each month contributes ~1 effective observation *per regime* (all 500 signals agree during broad trends — [correlated bets, one bet](/guides/evaluate-trading-backtest/#4-how-many-independent-bets-is-this-really)); you carry structural beta most of the time; the failure mode is whipsaw at trend turns. Same input, different risk factory. The evidence asymmetry follows: cross-sectional claims accumulate effective sample size fast (N assets × T periods, though cross-correlated), while time-series claims lean on far fewer independent regime episodes — a time-series result on 30 years of one index may rest on a dozen genuine trend cycles. Hold the two to [different significance standards](/guides/statistical-vs-economic-significance/); most published time-series evidence is thinner than its t-stats look. ## The checklist when a paper claims "predictability" - ☐ **Which claim is it?** Long-short-ranked = cross-sectional; own-history-conditioned = time-series. If the paper mixes constructions across tables, evaluate each separately. - ☐ **Effective sample size computed accordingly**: cluster cross-sectional stats by period; count time-series stats in regime episodes, not months. - ☐ **Beta accounting**: time-series strategies must report performance *against their embedded market exposure* — a long-biased trend rule "predicting" equities in a bull sample is [beta with a story](/guides/alpha-beta-alternative-risk-premia/). - ☐ **Crash mode identified**: skew and worst-drawdown by construction — reversal spikes (cross-sectional) vs whipsaw regimes (time-series). A paper silent on its construction's known failure mode hasn't met its own literature. - ☐ **Costs match the construction**: cross-sectional needs shorting and [borrow costs](/guides/transaction-costs-slippage-market-impact/); time-series needs turnover-at-regime-turns accounting. - ☐ **Breadth honesty**: cross-sectional edges scale with breadth (many independent names); time-series edges scale with independent *markets* — 50 correlated equity indices are not 50 bets. ## Questions to ask while reading 1. If I hedged the market factor out of this result, what remains? 2. How many *independent* episodes (not observations) support the claim? 3. Does the paper test the same signal in the *other* construction — and if the effect exists in only one, does the mechanism story explain why? 4. For ML papers: is the model predicting ranks or levels? [Loss-function choice](/guides/evaluate-ml-trading-paper/) silently picks a claim type. ## What our scoring captures — and misses Our [rigor scoring](/score-guide/) evaluates whether a paper's evidence supports *its* claim, but the claim-type distinction itself is carried in the paper's construction — and our archive shows the genres cluster: cross-sectional work dominates [factor investing](/topics/factor-investing/), time-series work runs through the trend and [volatility](/topics/volatility/) literatures, and ML papers switch constructions between tables more than any other genre. When browsing hubs, read the portfolio-construction sentence before the results table; it tells you which evidence standard to load. --- **Companions:** [Momentum's two families →](/guides/cross-sectional-vs-time-series-momentum/) · [Evaluating factor papers →](/guides/evaluate-factor-investing-paper/) · [Sample-size honesty →](/guides/evaluate-trading-backtest/) ---------------------------------------------------------------------- # Event-Driven Trading Research: Event Definitions, Leakage, and Execution URL: https://thequant.space/guides/event-driven-research/ Section: Guides Date: 2026-09-07 Description: The methodology of event studies for trading: defining events without hindsight, the announcement-timestamp problem, abnormal-return construction, and the execution window where paper edges vanish. **Not an implementation claim.** Event studies are among the most [leakage-prone](/guides/look-ahead-bias-point-in-time-data/) research designs in finance; this guide is the methodology for reading and building them without the classic self-deceptions. **Definition.** Event-driven research measures returns conditional on discrete occurrences — earnings, index changes, mergers, ratings actions, [resolutions in prediction markets](/guides/prediction-market-backtesting/). The design's power (clean identification around a timestamp) is exactly its fragility: everything depends on the event's *definition*, its *timestamp*, and the *window's tradability*, and each hides a failure mode. ## Problem 1: Defining events without hindsight An "event" must be identifiable **at the time, by the rule stated**. The violations are systematic: **conditioning on outcomes** ("merger announcements that completed" — completion is future information; the tradeable set includes the deals that broke, [the survivorship of events](/guides/survivorship-bias/)); **retroactive taxonomies** ("major" earnings surprises defined by thresholds fit on the full sample — [normalization leakage](/guides/look-ahead-bias-point-in-time-data/#6-normalization-and-statistics-computed-on-the-full-sample) in event clothing); and **database backfill** (event databases add, correct, and reclassify events after the fact — the point-in-time event list differs from today's download, and vendors rarely preserve the vintage). The discipline: write the event rule as an algorithm over [point-in-time data](/guides/look-ahead-bias-point-in-time-data/#the-point-in-time-discipline-positively-stated), then ask what its live false-positive rate would have been. ## Problem 2: The timestamp is the methodology Event *day* is not event *time*: earnings drop pre-open or post-close; news bodies timestamp distribution, not information creation; regulatory filings have acceptance vs dissemination times. Being wrong by hours flips which bar "reacts" — and a "predictive" result that evaporates under [correct timestamps is the genre's signature failure](/guides/auditing-a-monte-carlo/#11-verify-field-semantics-against-a-finer-dataset). Worse is **pre-event drift**: information leaks (analyst chatter, [order flow](/guides/limit-order-book-imbalance/)) mean the announcement window captures only the residual; a strategy "trading the event" competes with everyone who traded the leak. Papers must state the timestamp source and its precision — and results at daily resolution for intraday events deserve default suspicion. ## Problem 3: Abnormal returns and the clustered-events trap Event returns need a counterfactual: subtract expected returns via a factor model over the window. Two standing traps: **model choice moves results** for long windows (a 60-day "drift" is exquisitely sensitive to the [benchmark's factor set](/guides/alpha-beta-alternative-risk-premia/#the-sorting-regression)); and **event clustering breaks independence** — earnings seasons, sector-wide ratings waves, and index reconstitutions land together, so [N events ≠ N observations](/guides/evaluate-trading-backtest/#4-how-many-independent-bets-is-this-really), and cross-sectional t-stats on clustered events overstate wildly. Cluster by calendar date at minimum; the honest papers do. ## Problem 4: The execution window The paper measures close-to-close; the information arrived at 4:05pm. What's actually available is the **overnight gap plus the tradable session** — and event returns concentrate in the gap you couldn't trade. The checklist: does the entry price *postdate* the information's availability (open-after-announcement, not close-before)? Are [event-window spreads](/guides/transaction-costs-slippage-market-impact/#6-costs-and-slippage-are-convex-in-urgency) used, not calm-market ones — liquidity around events is exactly when spreads triple? For merger-style strategies: is the borrow available, and what does the [deal-break tail](/guides/high-sharpe-ratio-not-investable/#3-volatility-isnt-risk--skew-is) cost? Post-announcement-drift claims that survive open-entry, event-spread costing are rare — and the ones that do are the genre's [real findings](/topics/factor-investing/). ## Common ways event strategies fail in production - **The event feed differs from the research database**: live detection latency, misclassifications, and coverage gaps mean the production event stream is a noisier, slower version of the backtest's — [a data-vendor question](/guides/market-data-vendors/) papers never address. - **The gap captures the edge**: live P&L concentrates in fills you can't get; realized returns are the paper's minus the gap. - **Crowding at the timestamp**: everyone's algo reacts to the same feed; the [microstructure around events](/topics/market-microstructure/) is a queue race the backtest didn't model. - **Event-rule drift**: the definition that backtested cleanly gets "improved" in production — each improvement [a new trial](/guides/p-hacking-financial-research/). ## Questions to ask an event paper 1. Was the event set constructible, as defined, on the event date? 2. What timestamp source, at what precision — and do results survive open-entry? 3. How is clustering handled in the standard errors? 4. What are costs *in the event window*, and does the edge survive them? ## What our scoring captures Event studies score across the full [rigor range](/score-guide/), and the discriminators are exactly this guide's four problems — timestamp discipline and clustered-error handling are visible in a paper's text and scored accordingly. The [NLP/LLM hub](/topics/nlp-llm/) is the genre's current frontier (text-defined events), where the [training-cutoff leak](/guides/look-ahead-bias-point-in-time-data/#7-llm-training-data-leakage--the-newest-member) stacks on top of everything above; read that hub with this guide open. --- **The leakage taxonomy:** [look-ahead & PIT →](/guides/look-ahead-bias-point-in-time-data/) · **The window's costs:** [transaction costs →](/guides/transaction-costs-slippage-market-impact/) · **The text frontier:** [alternative-data map →](/guides/map-alternative-data/) ---------------------------------------------------------------------- # Free vs Paid Market Data: What Changes in Real Research URL: https://thequant.space/guides/free-vs-paid-market-data/ Section: Guides Date: 2026-09-07 Description: What separates free market data from paid: the six quality dimensions that matter for research, where free data is genuinely sufficient, and the failure modes that only surface after you've built on it. The question isn't whether free data is "good enough" in general — it's *which properties your research consumes*. Free and paid data differ on six specific dimensions, and a study that doesn't touch a dimension doesn't pay for skipping it. This guide maps the dimensions to research types; the vendor landscape itself is in the [market-data guide](/guides/market-data-vendors/). ## The six dimensions that actually separate the tiers **1. The dead.** The sharpest line: free sources overwhelmingly serve *current* listings — [survivorship bias built-in](/guides/survivorship-bias/#the-five-dead-tickers-test). Any cross-sectional equity strategy on free data starts 1–4%/yr optimistic. Crypto partially escapes (exchange APIs keep delisted pairs' history — sometimes) which is one reason [crypto research reproduces best](/guides/find-datasets-quant-papers/#the-standard-dataset-zoo). **2. Point-in-time integrity.** Free fundamentals are *current-vintage*: restated, backfilled, [look-ahead machines](/guides/look-ahead-bias-point-in-time-data/#1-restated-fundamentals--the-classic). There is essentially no free PIT fundamentals source; this dimension alone gates fundamental-signal research to paid tiers. **3. Corporate-action fidelity.** Free price series are usually pre-adjusted with undocumented policies — [dividend handling unknown, adjustment vintage unknown](/guides/corporate-actions-adjusted-prices/). Fine for charts; corrosive for total-return calculations and anything touching the [short leg](/guides/what-makes-a-factor-tradable/#gate-3-the-short-side--half-the-premium-most-of-the-problems). **4. Timestamps and completeness.** Free intraday data (where it exists) carries vague [timestamp semantics](/guides/look-ahead-bias-point-in-time-data/#5-timestamp-semantics), silent gaps, and no SLA on either. The failure mode is quiet: your pipeline interpolates the gap, and the [staleness detector you didn't build](/guides/notebook-to-production/#5-logs-metrics-and-the-alert-that-actually-matters) never fires. **5. Licensing.** Free tiers prohibit redistribution and often commercial use; "free for research" ends where [managing outside money begins](/guides/market-data-vendors/#5-what-does-the-license-actually-permit). Paid tiers sell *rights*, not just bytes — the dimension that matters precisely when success arrives. **6. Rate limits and history depth.** Free APIs cap requests and truncate history — a universe-wide daily download that takes three weeks through a rate limiter isn't free; it costs three weeks. ## Where free data genuinely suffices Honest cases for $0: **methodology development** (building your [backtest harness](/guides/quant-research-pipeline/#4-backtest--a-harness-not-a-script), testing pipeline plumbing — the data's biases don't matter because the results aren't claims); **index/futures-level strategies** (major-index daily series are commoditized and survivorship-immune at the index level — much of the [time-series momentum](/guides/cross-sectional-vs-time-series-momentum/) and [vol-targeting](/guides/volatility-targeting/) literature is testable this way); **crypto microstructure** (exchange APIs give real tick/book data free — [venue-quality caveats](/guides/microstructure-noise-realized-volatility/#common-ways-this-fails-in-production) applied); **factor research via pre-built libraries** (Ken French's portfolios are survivorship-handled *for you* — analysis on top of them inherits their construction, [stated and known](/guides/find-datasets-quant-papers/)); and **learning**, where the tuition is the point. The pattern: free data works when *someone else already paid the quality costs* (index providers, French library) or when the asset class's structure makes the dimensions moot. ## The upgrade triggers Move to paid the moment any of these becomes true: your universe includes individual equities cross-sectionally (dimension 1); your signals touch fundamentals (dimension 2); you're computing total returns or trading shorts (3); intraday timing enters the strategy (4); money — yours at scale, or anyone else's — depends on it (5); or re-downloading history has become a project (6). In [stack terms](/guides/quant-research-stack/#layer-1-market-data--own-your-pipeline-rent-the-feed): the $30–100/mo tier removes dimensions 4 and 6; dimensions 1–3 are what the serious-individual tier is *for*. ## Common ways free-data research fails late - **The result that was the bias**: months into a small-cap anomaly before the [five-dead-tickers test](/guides/survivorship-bias/#the-five-dead-tickers-test) reveals the universe was survivors-only. - **Silent API drift**: free endpoints change schemas and adjustment policies without notice; your [raw-zone archive](/guides/quant-research-stack/) of a moving target is an archive of inconsistencies. - **The republication trap**: a published result whose data can't be licensed for the follow-up. ## Questions to ask before building on any free source 1. Which of the six dimensions does my research consume — and does this source claim them? 2. Can I run the dead-tickers and [restatement spot-checks](/guides/market-data-vendors/#2-is-fundamental-data-point-in-time)? 3. What does the license permit at my ambition's end-state, not its start? ## What our scoring captures Papers on free data aren't penalized as such — [the rigor axis](/score-guide/) scores whether conclusions respect the data's limits. What the axis reliably catches is the mismatch: cross-sectional equity claims on survivors-only sources, fundamental signals with no PIT statement. The [data section's five-minute read](/guides/how-to-read-quant-finance-papers/#the-10-minute-pass-in-the-right-order) is where you catch it yourself. --- **The vendor landscape:** [market-data guide →](/guides/market-data-vendors/) · **What each data type answers:** [OHLCV to order book →](/guides/market-data-types/) · **The architecture:** [PIT data design →](/guides/point-in-time-data-architecture/) ---------------------------------------------------------------------- # From Notebook to Production: A Minimal Automated-Strategy Operating Stack (2026) URL: https://thequant.space/guides/notebook-to-production/ Section: Guides Date: 2026-09-06 Description: The minimal operating stack for running an automated trading strategy: one VPS, Docker Compose, Postgres, cron done right, secrets, logs, dead-man alerting, backups, and incident response — with working configs. Your backtest works. Your notebook produces signals. Now it has to run at 9:31 every morning without you — through reboots, API outages, fat-fingered configs, and the Tuesday you're on a plane. The gap between those two states is not more strategy code; it's an **operating stack**, and the minimal version is smaller than most engineers assume: one VPS, five containers, one database, and about a day of setup. This is the *how* companion to our [production-readiness checklist](/guides/production-checklist/) — that guide says what must be true; this one shows the smallest stack that makes it true. Scope: a solo operator or small team running daily-to-intraday strategies through a broker/exchange API. HFT has different answers at every layer. ## The stack at a glance | Layer | Minimal choice | Upgrade when | |---|---|---| | Host | One $10–25/mo VPS (Hetzner/DO/Vultr class) | Second region only when a real DR plan exists | | Runtime | Docker Compose, restart policies | Kubernetes: almost never for this workload | | State | Postgres (add Timescale for intraday bars) | Managed Postgres when ops time > $30/mo of value | | Scheduling | cron or systemd timers + run-locking | Airflow/Prefect only with many interdependent jobs | | Secrets | `.env` at 600 + Docker secrets | SOPS/age or a vault when a second human joins | | Logs | JSON to stdout → journald/Docker, 30-day rotation | Loki/hosted only when you actually grep across weeks | | Metrics/alerts | Dead-man heartbeat (Healthchecks/Uptime Kuma) + ntfy to phone | Prometheus + Grafana at 2+ live strategies | | Backups | Nightly `pg_dump` → offsite object storage + restore drill | Streaming replication when RPO of 24h hurts | The theme: **every component is boring on purpose.** Your strategy is the experiment; the stack around it must not be. ## 1. One VPS, containers as the unit of operation Run everything as Docker Compose services on a single small VPS. Containers earn their place here for three unglamorous reasons: the image pins your Python and dependency versions permanently (the [checklist's](/guides/production-checklist/#gate-2--code) reproducible-environment box), `restart: unless-stopped` gives you crash recovery without writing supervisor configs, and a deploy becomes `git pull && docker compose up -d --build` — which means rollback is `git checkout ` and the same command. The skeleton that carries the whole guide: ```yaml # docker-compose.yml services: strategy: build: . restart: unless-stopped env_file: .env # secrets live here, mode 600, not in git depends_on: [db] logging: driver: json-file options: { max-size: "20m", max-file: "5" } watchdog: # kill switch: SEPARATE container, minimal deps build: ./watchdog restart: unless-stopped env_file: .env.watchdog # its own broker credentials, read-only where possible db: image: timescale/timescaledb:latest-pg16 restart: unless-stopped environment: POSTGRES_PASSWORD_FILE: /run/secrets/pg_pass secrets: [pg_pass] volumes: ["pgdata:/var/lib/postgresql/data"] secrets: pg_pass: { file: ./secrets/pg_pass } volumes: pgdata: ``` Notes that save future pain: pin image tags in real life (`timescaledb:2.17.2-pg16`, not `latest`); the watchdog deliberately does **not** share code or an image with the strategy — an independent process that flattens positions on heartbeat loss is the highest-value component in the stack and it must not inherit the strategy's bugs; and put the VPS behind key-only SSH with a firewall allowing 22 and nothing else inbound — this stack needs no open ports. ## 2. Postgres is the source of truth, not your process memory The single biggest architectural decision: **the strategy process must be disposable.** Kill it at any moment and a restart rebuilds state from two places only — the broker API and the database. Anything held in memory or in loose files is state you'll lose at the worst time. Minimum schema, four tables: ```sql CREATE TABLE runs (id bigserial PRIMARY KEY, started_at timestamptz DEFAULT now(), finished_at timestamptz, status text, detail jsonb); CREATE TABLE decisions (id bigserial PRIMARY KEY, run_id bigint REFERENCES runs, ts timestamptz DEFAULT now(), symbol text, action text, inputs jsonb, reason text); -- the decision log CREATE TABLE orders (id bigserial PRIMARY KEY, client_order_id text UNIQUE, decision_id bigint REFERENCES decisions, ts timestamptz, symbol text, side text, qty numeric, status text, broker_id text); CREATE TABLE fills (id bigserial PRIMARY KEY, order_id bigint REFERENCES orders, ts timestamptz, qty numeric, price numeric, fee numeric); ``` `client_order_id UNIQUE` is doing real work: it's the database-enforced half of order idempotency — on any uncertainty, reconcile against the broker before re-sending, and the constraint makes accidental double-submission a loud error instead of a double position. The `decisions` table with full `inputs` jsonb is what lets you answer "why did it trade?" months later. **Do you need Timescale?** For a daily strategy storing signals and orders: plain Postgres, no. The Timescale extension earns its place when you're persisting intraday bars or tick captures next to your operational tables — hypertables + compression keep that from eating the disk. Either way this pairs with the research-side storage question covered in the [tick-database guide](/guides/tick-data-databases/). And honestly: a single-process daily strategy *can* run on SQLite — you give up concurrent access for the watchdog and painless backups, which is exactly why we don't recommend it once real money flows. ## 3. Scheduling: cron is fine, double-runs are not For a handful of independent jobs (pre-market prep, signal run, EOD reconciliation), **cron or systemd timers are the correct tool** — Airflow at this scale is an incident generator with a UI. What cron does not give you is protection from the two classic failures: overlapping runs and silent no-runs. Overlap protection in one line with a Postgres advisory lock (works across containers, unlike `flock` on a file): ```python got_lock = db.execute("SELECT pg_try_advisory_lock(42)").scalar() if not got_lock: log.warning("previous run still holds the lock; exiting"); sys.exit(0) ``` Silent no-runs are solved by the dead-man pattern in §5 — the job's *absence* pages you, cron's own logs never will. And record every invocation as a `runs` row first thing; "when did this last actually execute" should be a SQL query, not an archaeology dig through syslog. One non-negotiable from the [checklist](/guides/production-checklist/#gate-2--code): the host runs NTP and every timestamp in the system is UTC. Exchange-calendar logic (holidays, half-days, DST transitions against a US-market clock) lives in one tested module, not scattered through crontabs. ## 4. Secrets: boring, strict, rotatable The minimal discipline that survives audits and mistakes alike: - Secrets live in `.env` files owned by the deploy user, mode `600`, listed in `.gitignore`, and injected via `env_file:`/Docker secrets — never baked into images, never in compose files, never in code. - **Separate credentials per component.** The strategy gets trade-scoped API keys; the watchdog gets its own; backups get read-only DB access. When one leaks, you rotate one. - Write the rotation procedure down *now* (where each key is issued, where it's referenced, how to verify the old one is dead). Rotation you've never rehearsed is downtime with extra steps. - If your broker supports IP allowlisting, pin the API keys to the VPS's address — it converts a leaked key from an emergency into a nuisance. Upgrade to SOPS/age-encrypted secrets in git, or a proper vault, when a second person needs access — not before. ## 5. Logs, metrics, and the alert that actually matters **Logs:** emit structured JSON to stdout and let Docker/journald handle files and rotation (the `logging:` block above caps it at 100MB). Two streams matter and they are not the same thing: the *application log* (exceptions, API latencies, retries) and the *decision log* (§2, in Postgres, queryable forever). Ship logs to Loki/a hosted service only when you find yourself actually needing cross-week searches. **The one alert that matters** is the dead-man switch: the strategy pings after each successful cycle; the monitor pages you when the ping *doesn't* arrive. This inverts the failure mode — crash, hang, expired credentials, dead VPS, and broken cron all collapse into the same phone notification. ```bash # last line of a successful run curl -fsS -m 10 "https://hc-ping.com/" # Healthchecks.io (hosted, free tier) # or self-hosted Uptime Kuma's push monitor — same pattern ``` Route pages through **ntfy** (self-hostable) or Pushover/Telegram — anything that vibrates a phone. Email is where alerts go to be read on Thursday. The watchdog container consumes the same heartbeats from Postgres and acts (flatten, halt) instead of merely notifying — alerting tells *you*, the kill switch protects the *account*, and you need both because you sleep. Full Prometheus + Grafana is genuinely worth it at two or more live strategies or when you start caring about API-latency trends — as a starting point it's ceremony. A daily P&L-and-positions summary pushed to the same phone channel doubles as your reconciliation nudge. ## 6. Backups you have restored at least once The state worth money is small: the Postgres volume and your config. That makes the minimal plan cheap and honest: ```bash # /etc/cron.d/backup — 02:15 UTC nightly 15 2 * * * deploy docker exec db pg_dump -U postgres -Fc trading \ | age -r $BACKUP_PUBKEY \ | rclone rcat r2:tq-backups/pg/trading-$(date -u +\%F).dump.age \ && curl -fsS -m 10 https://hc-ping.com/ ``` Encrypted with `age`, shipped **off the VPS** to object storage (R2/S3/B2 — cents per month at this size), with its own dead-man check so a silently failing backup pages you like a silently failing strategy. Keep 30 daily + 12 monthly via a lifecycle rule. Code and config need no separate backup — they're in git; the `.env` secrets go in your password manager, not in the backup stream. The part everyone skips: **quarterly, restore the latest dump to a scratch container and run one query against it.** The checklist's kill-switch drill and this restore drill are the same idea — an untested recovery path is a decoration. RPO here is 24 hours of *operational metadata* (fills are recoverable from the broker); if that ever becomes unacceptable, that's the trigger for WAL archiving or streaming replication, not before. ## 7. Incident response for a team of one You will have incidents; the difference between a bad hour and a bad quarter is whether the middle-of-the-night decisions were made in advance: - **A one-page runbook** in the repo: how to halt trading, how to flatten manually at the broker, how to restart from nothing on a fresh VPS (you tested this when you set up backups), who to call at the broker. - **An incident log** (a markdown file is fine): timestamp, symptom, cause, fix, follow-up. Its compounding value is the pattern — the third entry about the same component is a redesign order, per the [checklist](/guides/production-checklist/#gate-6--go-live-protocol). - **Pre-decided halt conditions.** Data feed stale > N minutes → halt new orders. Reconciliation mismatch → halt and page. Drawdown breach → watchdog flattens without asking you. The strategy resumes only by explicit human action — auto-resume after an unexplained halt is how one incident becomes two. ## Use this stack if / don't use it if **Use this if:** you run 1–3 strategies at minute-to-daily horizons through a broker API, you are the ops team, and your monthly infrastructure budget has one digit before the comma. This stack — VPS, Compose, Postgres, cron with locks, dead-man alerts, offsite dumps — covers every box in Gates 3–5 of the checklist for roughly **$15–35/month** all-in. **Don't use this if:** latency is your edge (you need proximity hosting and a different article), you have a team (add SOPS, CI deploys, and real on-call rotation), regulation requires audit infrastructure beyond a decision log, or you haven't passed [Gate 1 validation](/guides/production-checklist/#gate-1--validation-before-writing-any-production-code) yet — production infrastructure for an unvalidated strategy is a way to lose money more reliably. *Tool mentions are editorial, dated September 2026; no affiliate links or sponsored placements on this page — if that changes it will be disclosed inline. Verify pricing before committing.* --- **The checklist this implements:** [Production-readiness checklist →](/guides/production-checklist/) ([download the .md](/downloads/production-readiness-checklist.md)) · **The research side:** [A practical quant research stack →](/guides/quant-research-stack/) ---------------------------------------------------------------------- # How to Backtest Prediction-Market Strategies Without Fooling Yourself (2026) URL: https://thequant.space/guides/prediction-market-backtesting/ Section: Guides Date: 2026-09-06 Description: Prediction-market backtests fail differently than equity backtests: resolution leakage, phantom liquidity, and survivorship in resolved markets. A methodology that survives contact with live books. Prediction markets look like the easiest possible backtesting target: prices are probabilities, outcomes are binary, resolution is objective. That surface simplicity is exactly why most published prediction-market "edges" evaporate on contact with a live order book. The failure modes here are different from equities, and equity-quant reflexes miss them. This guide is the methodology we'd hold any strategy to before real capital touches a CLOB — written from the perspective of someone who has watched paper edges die in production. ## Why prediction markets break equity intuitions Four structural differences drive everything below: 1. **Terminal, binary payoffs.** Every position resolves to 0 or 1 by a known-ish date. There's no "hold and be right eventually" — timing and capital lockup are part of the bet. 2. **Prices *are* probabilities**, so calibration analysis — not just P&L — is available and mandatory. 3. **Books are thin.** Outside headline politics/sports markets, visible depth is often a few hundred to a few thousand dollars within a cent of mid. Any backtest that fills meaningful size at mid is fiction. 4. **Markets are born and die constantly**, and the historical record you can easily download is dominated by markets that *resolved cleanly* — a survivorship filter nobody advertises. ## The five biases, in the order they'll get you ### 1. Resolution leakage (look-ahead's nastier cousin) The classic error: your feature set knows things that were only knowable near resolution. It sneaks in through innocuous routes — using a market's *final* metadata (titles and descriptions get edited), computing features over a window that overlaps the resolution announcement, or training an NLP model on news text whose corpus extends past the event. The discipline: every feature must be computable from a snapshot taken strictly before the trade timestamp, and market metadata must be the *creation-time* version. If you can't reconstruct point-in-time metadata, exclude the feature. ### 2. Phantom liquidity Prediction-market backtests are usually built on trade prints or last-price series, then evaluated as if you could trade at those prices in your size. In thin books this overstates capacity by an order of magnitude. Minimum honest fill model: - Cross the spread: buys lift the ask, sells hit the bid — never fill at mid or last. - Cap fill size at a fraction (say a third) of visible depth at that level, or walk the book if you have depth data. - Apply the platform's actual fee schedule (taker fees, settlement fees, and on-chain venues, gas) per fill, not per backtest. **Record your own books.** Most venues do not provide historical order-book depth; trade prints are not depth. A cron job snapshotting the CLOB for your candidate markets every few minutes costs almost nothing and is the single highest-value data asset a prediction-market trader can own. Start recording *before* you need the history — nobody sells it to you later. ### 3. Survivorship and the void-market problem Resolved-market datasets silently exclude or mishandle markets that were voided, resolved "N/A," got disputed, or had their terms clarified mid-life. Strategies harvesting "cheap longshots" or "expensive favorites" are especially distorted: disputed and voided markets cluster exactly in the weird tails where those strategies live. Your universe at each point in time must be "markets tradeable at time t," including the ones that later resolved ugly — and dispute/void outcomes must hit the P&L (usually as a refund minus fees and a lot of locked capital time). ### 4. Correlated event clusters Fifty markets on one election are one bet wearing fifty costumes. Naive per-market analysis reports n=50 independent wins when the true sample size is 1. Cluster by underlying event before computing any statistics; a strategy with 400 trades across 12 event clusters has 12 data points for significance purposes. This is the prediction-market version of the multiple-testing problem that wrecks [factor research](/topics/factor-investing/). ### 5. Capital lockup and the funding illusion A 4% edge that locks capital for nine months is a worse trade than a 1% edge that resolves weekly. Backtests that report return-per-trade without annualizing per unit of locked capital systematically favor long-dated markets. Track capital-days per position and report **return on committed capital per unit time** — it reorders most strategy leaderboards. ## Calibration before P&L Before any profit claim, run the sanity layer that binary markets uniquely allow: - **Reliability curve**: bucket your model's predicted probabilities, plot realized frequencies. A model that says 70% should be right ~70% of the time. - **Brier score / log loss** against two baselines: the market price itself, and a naive "always the market price minus fees" model. If your model doesn't beat the *market's own* calibration, your edge is fee-negative noise. - **Edge decomposition**: is the P&L from probability estimation, from liquidity provision (earning spread), or from timing flows? These have completely different capacity and decay profiles, and mixing them in one backtest number hides which one you actually have. ## The walk-forward protocol 1. **Universe construction in event time**: at each historical date, list markets that existed, were tradeable, and met your liquidity floor *using only that date's information*. 2. **Strict temporal splits** for any learned model, with embargo periods around cluster boundaries (train on 2023–24 politics, test on 2025+ — never interleave markets from one event across splits). 3. **Fee-and-spread-first reporting**: publish the gross-to-net waterfall. In our experience the spread+fee haircut on thin markets routinely consumes 60–100% of apparent gross edge. 4. **Paper trading against live books** for 4–8 weeks minimum, with your real execution code, comparing realized fills to backtest-assumed fills. The fill-quality gap is your ongoing reality-check metric, not a one-time gate. 5. **Tiny-size live period** before scale, per the [research-to-production checklist](/guides/production-checklist/). ## Tooling that makes this tractable - Venue CLOB/data APIs for live books and trade prints; your own snapshot archive for depth history (Parquet + DuckDB handles years of it — see the [stack guide](/guides/quant-research-stack/)). - Resolution-source archives (the oracle/UMA record, official statistics pages) with timestamps, so resolution timing itself is a modeled variable. - A position/risk monitor with per-event exposure caps — the correlated-cluster bias exists in live trading too, not just backtests. *We're packaging this methodology as a runnable backtest template (notebook + fill model + calibration suite). Subscribers hear first when it ships — signup below.* --- **Related:** [Crypto & DeFi research hub →](/topics/crypto-defi/) · [Production checklist →](/guides/production-checklist/) ---------------------------------------------------------------------- # How to Build a Quant Research Pipeline: From Idea to Evidence, Repeatably URL: https://thequant.space/guides/quant-research-pipeline/ Section: Guides Date: 2026-09-07 Description: The seven-stage quant research pipeline — idea intake, data, features, backtest, validation, paper trading, production — with the artifact each stage must produce and the gate it must pass. A notebook answers a question once. A **pipeline** answers questions repeatably — with the biases handled in one place, the trials counted, and every promising result already wearing the harness it needs for production. This guide assembles the whole system the other guides cover piecewise: seven stages, and for each one, the *artifact* it must produce and the *gate* it must pass. It is also the map of where infrastructure spending actually pays — which is why it links outward to the stack, database, and tooling guides. ## The seven stages ```text idea → data → features → backtest → validation → paper trading → production │ │ │ │ │ │ │ log raw zone store result row verdict fill stats operating stack ``` ### 1. Idea intake — the experiment log Every idea gets a row *before* any code: hypothesis, motivation (a paper? a [topic hub](/topics/)? an anomaly in live fills?), and its predeclared success criteria. This log is not bureaucracy — it is the **trial counter** that makes your [deflated Sharpe ratio](/guides/deflated-sharpe-ratio/#using-it-honestly-without-ceremony) computable and your [significance claims](/guides/statistical-vs-economic-significance/) honest. Abandoned ideas stay in the log; they count. **Artifact:** one row per idea, append-only. **Gate:** written success criteria before data is touched. ### 2. Data — one clean layer, biases handled once The pipeline's foundation is a data layer where the classic failures are fixed *centrally*, so no individual study can reintroduce them: [survivorship-clean universes](/guides/survivorship-bias/) reconstructed as-of-date, [point-in-time semantics](/guides/look-ahead-bias-point-in-time-data/) with knowable-timestamps on every field, raw vendor payloads archived append-only, and verified field semantics ([cross-checked against finer data](/guides/auditing-a-monte-carlo/#11-verify-field-semantics-against-a-finer-dataset)). Vendor selection criteria are in the [market-data guide](/guides/market-data-vendors/); storage architecture — Parquet + DuckDB first, [a real database when it hurts](/guides/tick-data-databases/) — in the [stack guide](/guides/quant-research-stack/). **Artifact:** the raw zone + a documented loader every study must use. **Gate:** the five-dead-tickers test and a PIT spot-check, rerun whenever a vendor changes. ### 3. Features — computed once, leak-proof by construction A shared feature library beats per-notebook feature code for one decisive reason: **point-in-time correctness gets engineered once**, with expanding-window transforms and announcement-lag handling built in, instead of being re-risked in every study ([leak taxonomy §6](/guides/look-ahead-bias-point-in-time-data/#6-normalization-and-statistics-computed-on-the-full-sample)). Version the definitions; a changed feature invalidates cached results downstream. This library is your actual IP — treat it like the production code it will become. **Artifact:** versioned feature definitions + point-in-time feature values. **Gate:** every feature passes a one-bar-lag sanity check on a known case. ### 4. Backtest — a harness, not a script One backtesting harness, shared by every study, embedding the house rules: next-bar execution, [asset-class cost priors](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges) with the 1×/2×/3× sensitivity curve computed by default, turnover and drawdown always reported, and results written to a **results database** — not screenshots — keyed to the experiment-log row. The engine choice matters less than the shared-ness ([stack guide's take](/guides/quant-research-stack/#layer-4-backtesting--the-layer-where-honesty-lives)); what compounds is that every result in your shop is comparable and countable. **Artifact:** a result row (metrics, config hash, cost curve) per run. **Gate:** the harness reproduces a known-answer strategy to machine precision ([the strongest validation pattern](/guides/auditing-a-monte-carlo/#61-test-invariants-not-outputs)). ### 5. Validation — the adversarial stage Everything before this stage generates candidates; this stage kills them. The full arsenal is already documented: the [eight backtest questions](/guides/evaluate-trading-backtest/), [purged walk-forward with an evaluate-once holdout](/guides/walk-forward-out-of-sample-testing/), parameter-plateau checks, and the [DSR hurdle](/guides/deflated-sharpe-ratio/) using the experiment log's honest N. For external ideas, this stage is the [reproduction checklist](/guides/reproduce-quant-research/). The output is a written verdict against the stage-1 criteria — by someone wearing the adversary hat, even when both hats are yours. **Artifact:** a verdict memo. **Gate:** [production checklist Gate 1](/guides/production-checklist/#gate-1--validation-before-writing-any-production-code), boxes actually ticked. ### 6. Paper trading — the assumptions meet the market The strategy runs through its *real* execution path against live markets at zero size, measuring the one thing no backtest can: the **fill-quality gap** between simulated and actual executions ([the ongoing cost measurement](/guides/transaction-costs-slippage-market-impact/#live-costs-are-the-ongoing-measurement)). For venues without good historical depth — [prediction markets](/guides/prediction-market-backtesting/#2-phantom-liquidity), long-tail crypto — this stage also begins recording the order books your next research cycle will need. **Artifact:** fill-gap statistics vs. backtest assumptions. **Gate:** gap within pre-set tolerance for 2–8 weeks. ### 7. Production — the operating stack The survivor gets the [minimal operating stack](/guides/notebook-to-production/): the same signal code (never a reimplementation), Postgres as source of truth, kill switch, dead-man alerting, daily reconciliation, tested backups. Live divergence stats flow back into the experiment log — closing the loop, because live performance is [the only out-of-sample test nobody can fake](/guides/walk-forward-out-of-sample-testing/#the-out-of-sample-tests-nobody-can-fake), and today's live anomalies are stage-1 rows for next quarter. **Artifact:** a deployed, monitored strategy. **Gate:** [Gates 2–6](/guides/production-checklist/) plus the go-live ramp. ## Build order for one person Don't build stages 1–7 as a project; grow them in the order that pays: **(1)** the experiment log — a table, an afternoon, permanent honesty gains; **(2)** the data layer's raw zone + loader ([stack guide layer 1–2](/guides/quant-research-stack/)); **(3)** the shared backtest harness with default cost curves; **(4)** the validation ritual as a written checklist; **(5–7)** only when a strategy earns them. Infrastructure bought ahead of a bottleneck is [the trap the stack guide warns about](/guides/quant-research-stack/#what-it-costs-september-2026-order-of-magnitude); infrastructure that makes trials countable and biases un-reintroducible pays from week one. Budget the compute side with the [cost calculator](/tools/compute-cost-calculator/). --- **The infrastructure map:** [research stack →](/guides/quant-research-stack/) · [databases →](/guides/tick-data-databases/) · [production stack →](/guides/notebook-to-production/) · **The evaluation arsenal:** [backtest questions →](/guides/evaluate-trading-backtest/) ---------------------------------------------------------------------- # How to Choose Which Papers to Replicate First URL: https://thequant.space/guides/choose-papers-to-replicate/ Section: Guides Date: 2026-09-07 Description: A prioritization framework for replication: the expected-information calculation, the five selection criteria, portfolio-of-replications thinking, and which paper types repay the effort. **The problem.** [Replication done right](/guides/reproduce-quant-research/) costs one to four weeks per paper. Your queue of candidates is effectively infinite. Choosing well is therefore a portfolio problem: maximize what you *learn per week* — about markets, about a literature's reliability, and about strategies you might run — rather than chasing the single most exciting result. ## The expected-information frame A replication's value = (probability the outcome changes what you do) × (size of that change) ÷ (weeks of effort). This immediately demotes two popular choices: **papers you're sure are right** (confirmation teaches little — replicate the canonical result only as [pipeline calibration](/guides/quant-research-pipeline/#4-backtest--a-harness-not-a-script), a different purpose) and **papers you're sure are wrong** (debunking a paper nobody was going to trade is entertainment). The information-rich zone is *genuine uncertainty about a result you'd act on* — a strategy you might run, a method you might adopt, a data source you might trust. ## The five selection criteria Score each candidate 0–2 on all five; replicate from the top: **1. Actionability (×2 weight).** Would a successful replication change your allocation, your [pipeline](/guides/quant-research-pipeline/), or your priors about a whole literature? Passing the [implementability tests](/guides/academic-vs-implementable/#the-five-implementability-tests) scores 2; pure curiosity scores 0. **2. Feasibility.** [Data obtainable at your tier](/guides/find-datasets-quant-papers/#the-access-tier-reality), compute within budget, methods within skill. A perfect candidate you can't finish is a 0 — and the most common planning error. Check [code availability](/guides/find-code-finance-papers/) here too: an official repo halves the cost, a good third-party repo helps *and* adds [audit obligations](/guides/find-code-finance-papers/#the-15-minute-audit-before-found-code-touches-your-research). **3. Informativeness of the *literature*, not just the paper.** Some papers are load-bearing: dozens of follow-ups inherit their construction ([FI-2010 benchmarks](/guides/map-lob-prediction/), a zoo factor's original sort, a widely-forked implementation). Replicating a load-bearing paper prices an entire branch of [a research map](/guides/map-ml-asset-pricing/) at once. Our maps and [hub rankings](/topics/) are built to spot these hubs-within-hubs. **4. Age sweet spot.** Too new (< ~1 year): no post-publication data, [decay unmeasurable](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer), and the authors' own robustness process may still be running. Too old with heavy citation: probably already replicated — *search for existing replications first* ([the critique search](/guides/search-quant-finance-literature/#search-tactics-that-pay-for-themselves)); reading five is cheaper than running one. The sweet spot: 2–6 years old, cited, unreplicated, with post-sample data now available — replication and [out-of-sample test](/guides/walk-forward-out-of-sample-testing/#the-out-of-sample-tests-nobody-can-fake) in one effort. **5. Suspicion asymmetry.** Prefer papers where your prior and the paper's claim *diverge* — high [rigor score](/score-guide/) but implausible magnitude, or modest claims from a famously careful group. Divergence is where updating happens. A [red-flag-free paper](/guides/evaluate-trading-backtest/#the-60-second-version) claiming SR 3 is exactly the profile that pays. ## Build a replication portfolio, not a pick Across a quarter, mix three types: one **calibration replication** (a settled result, to validate your [harness](/guides/quant-research-pipeline/#4-backtest--a-harness-not-a-script) and data — expected outcome: match); one **decision replication** (the strategy or method you might actually adopt — the actionability play); and one **audit replication** (a load-bearing or suspicious paper — the public-good play, and the source of the field's scarcest artifact: [published replication notes](/guides/reproduce-quant-research/#stage-6--document-make-your-reproduction-reproducible)). The mix hedges the portfolio: calibrations always pay something, decisions pay privately, audits pay reputationally — our [T-KAN review](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion) was an audit replication, and it repriced a result. ## Questions to ask before committing a month 1. What specific decision changes on success? On failure? (If the answers match, don't run it.) 2. Has someone already done this? (Twenty minutes of [critique-searching](/guides/search-quant-finance-literature/) first, always.) 3. What's the *minimum* version — one table, one market, one subperiod — that delivers 80% of the information? 4. Will I publish the notes either way? (Precommit; failed replications are the ones the field needs most.) ## How to use this site for the selection This is the workflow the archive was built for: filter a [topic hub](/topics/) by rigor, cross-reference the [research map](/guides/map-statistical-arbitrage/) for load-bearing branch positions, run the [ten-minute read](/guides/how-to-read-quant-finance-papers/) on candidates, score them on the five criteria — and when you finish a replication, leave the notes in the paper's discussion section, where the next selector will find them. --- **The execution manual:** [reproduction checklist →](/guides/reproduce-quant-research/) · **The triage upstream:** [how to read papers →](/guides/how-to-read-quant-finance-papers/) · [backtest evaluation →](/guides/evaluate-trading-backtest/) ---------------------------------------------------------------------- # How to Distinguish an Academic Contribution from an Implementable Idea URL: https://thequant.space/guides/academic-vs-implementable/ Section: Guides Date: 2026-09-07 Description: A framework for separating papers that advance knowledge from papers you can trade: the five implementability tests, why both kinds have value, and how the quadrant system encodes the distinction. **The distinction.** An **academic contribution** advances what the field knows: a cleaner identification, a new estimator, a documented regularity. An **implementable idea** survives contact with capital: obtainable data, executable trades, [costs that don't consume it](/guides/transaction-costs-slippage-market-impact/), capacity worth the effort. These overlap far less than readers assume, because the incentives point different directions — publication rewards novelty and statistical cleanliness; implementation rewards robustness and boring mechanics. Neither is superior; confusing them is how researchers waste quarters and traders fund [other people's noise](/guides/p-hacking-financial-research/). ## The five implementability tests Run any promising paper through these, in order of how fast they disqualify: **1. The data test.** Is the input data [buyable at your tier](/guides/find-datasets-quant-papers/#the-access-tier-reality), at production latency? A signal from quarterly institutional holdings data published with a 45-day lag is an academic instrument, not a trading input — unless the paper's *horizon* respects that lag. **2. The friction test.** Net the claimed edge against [realistic costs for its turnover and universe](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges). The academic convention of gross returns isn't dishonest — cost structures are reader-specific — but it means *you* run this test, always. Most cross-sectional results with monthly-or-faster rebalancing fail it. **3. The capacity test.** [At what size does this stop working](/guides/capacity-constraints/), and is that size worth your infrastructure? Effects in micro-caps and thin books can be simultaneously true, published, and irrelevant to any real portfolio — [investability's second link](/guides/high-sharpe-ratio-not-investable/#2-it-doesnt-scale). **4. The staleness test.** Published effects [decay](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer); the relevant question is post-publication, post-crowding performance. A 2015 paper's 1990–2014 backtest is an artifact of history; its 2015–2026 behavior is the [only test that matters](/guides/walk-forward-out-of-sample-testing/#the-out-of-sample-tests-nobody-can-fake). **5. The infrastructure test.** What does capturing this *operationally require* — tick data, colocation, borrow relationships, [a production stack](/guides/notebook-to-production/)? An edge whose infrastructure costs exceed its dollar alpha at your size fails implementability even if everything else passes. ## What to do with each kind **Pure academic contributions are still valuable to practitioners** — differently: a new estimator improves your [research pipeline](/guides/quant-research-pipeline/) even if its example application is untradeable; a documented regularity that fails the friction test may still *inform* position sizing or timing in strategies you already run; and identification-focused papers teach mechanisms, which is what [robustness stories](/guides/robustness-trading-strategy/#2-universe-robustness--the-transfer-test) are made of. The failure mode is only the category error: backtesting an instrument that was never a strategy. **Implementable ideas from papers are starting points, not endpoints**: the published version is the *most decayed, most crowded* version of the idea. Implementation work — better execution, universe extension, [regime conditioning](/guides/regime-dependence/#what-to-do-with-a-regime-dependent-edge) — is where the residual edge lives, which is why the [replication-first workflow](/guides/choose-papers-to-replicate/) beats the invent-first one for most independent researchers. ## The tells, at a glance | Signal | Leans academic | Leans implementable | |---|---|---| | Returns reported | Gross, value-weighted alphas | Net, with turnover and capacity | | Data | Institutional, historical, lagged | [Buyable, current](/guides/market-data-vendors/) | | Horizon | Monthly+ formation periods | Matches obtainable execution | | Robustness section | Statistical (t-stats under specs) | Operational ([costs, delays, sloppiness](/guides/robustness-trading-strategy/#5-implementation-robustness--the-sloppiness-test)) | | Authors | Pure academic | Practitioner co-authors, live-tested claims | ## Questions to ask while reading 1. If this is true, *who is paid to remove it* — and why haven't they? (No answer = probably [not implementable](/guides/evaluate-trading-backtest/); a good answer names the constraint you'd need to escape.) 2. What's the paper's implicit AUM, and what's yours? 3. Which of the five tests does the paper itself address? (Each one addressed is a rigor signal.) 4. Is the contribution the *result* or the *method*? Methods age better. ## How the quadrant system encodes this This distinction is half of why our [quadrant system](/score-guide/) exists: **Lab Rats** (high math, low rigor) is largely the academic-contribution quadrant; **Street Traders** (empirical, lighter theory) leans implementable; **Holy Grail** papers pass both readings. But the mapping is deliberately imperfect — empirical rigor measures *evidence quality*, and the five tests above measure *your* frictions, capacity, and infrastructure. The score answers "is this claim well-supported?"; only you can answer "is this claim mine to use?" --- **The bridge from idea to capital:** [replication first →](/guides/choose-papers-to-replicate/) · [the research pipeline →](/guides/quant-research-pipeline/) · [production readiness →](/guides/production-checklist/) ---------------------------------------------------------------------- # How to Evaluate a Factor-Investing Paper URL: https://thequant.space/guides/evaluate-factor-investing-paper/ Section: Guides Date: 2026-09-07 Description: A checklist for evaluating factor-investing and cross-sectional asset-pricing papers: multiple-testing hurdles, portfolio construction choices, cost realism, and the questions that expose a factor-zoo entry. **Definition.** A factor-investing paper claims that a characteristic — value, momentum, quality, or one of [several hundred published successors](/topics/factor-investing/) — predicts the cross-section of returns: sort assets by it, go long the top bucket and short the bottom, and earn a premium. The genre is the most published and least replicable in empirical finance, which makes the *generic* [backtest checks](/guides/evaluate-trading-backtest/) necessary but not sufficient. These are the factor-specific ones. ## The worked example: construction is the hidden factor The same signal can produce a t-stat of 1 or 4 depending on choices that rarely make the abstract. A characteristic sorted into deciles, equal-weighted, including micro-caps, rebalanced monthly, 1963–2000 might show 8%/yr of spread; the identical characteristic in quintiles, value-weighted, NYSE-breakpoints, ex-micro-caps often shows a third of that. Nothing was "wrong" — each cell in that grid is a separate trial, and the paper you're reading is the cell that worked. This is why the [multiple-testing hurdle](/guides/statistical-vs-economic-significance/#why-the-usual-significance-bar-is-broken-in-finance) for new factors sits near **t ≈ 3.0, not 2.0**. ## The factor-paper checklist - ☐ **Weighting and breakpoints**: value-weighted results with NYSE breakpoints reported? Equal-weighted deciles are a micro-cap bet wearing a factor costume — and micro-caps are where [survivorship](/guides/survivorship-bias/) and untradeable [spreads](/guides/transaction-costs-slippage-market-impact/) concentrate. - ☐ **The long and short legs separated**: many premia live entirely in the short leg — expensive to borrow, sometimes impossible. A "premium" you can only earn by shorting unborrowable names is a museum piece. - ☐ **Post-formation lag**: characteristic measured with [announcement-date information only](/guides/look-ahead-bias-point-in-time-data/#2-reporting-lag-compression)? Accounting signals used at fiscal period-end are leaking. - ☐ **Turnover and capacity stated**: monthly-rebalanced factors can turn over 100–400%/yr; multiply by realistic costs before believing net claims. - ☐ **Subperiods, including post-2003**: decimalization, and later the publication of the factor itself, mark regime lines. A factor alive only pre-2000 is history, not a strategy — the [decay literature](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer) is unambiguous. - ☐ **Spanning tests**: is the new factor priced *after* controlling for the standard set, or is it old exposure with a new name? ## Questions to ask while reading 1. How many characteristic/construction combinations does the paper's own appendix reveal were examined? 2. Would the result survive value-weighting minus micro-caps? (If the paper doesn't show it, assume no.) 3. What is the net-of-cost premium *at the fund size the authors imply*? 4. Does the factor exist in international samples the authors didn't optimize on? 5. If published five-plus years ago: what happened after publication? ## What our scoring captures — and misses The [rigor axis](/score-guide/) rewards exactly what this genre most often lacks: cost accounting, out-of-sample sections, and disclosed robustness grids, so high-rigor factor papers are disproportionately trustworthy. What the score *cannot* see is the unpublished search space — a clean-looking paper that quietly tried 200 constructions scores the same as one that pre-registered a single design. That correction is yours to apply, via the [deflated Sharpe ratio](/guides/deflated-sharpe-ratio/) with an honest N; our score is a floor on quality, not a ceiling on skepticism. --- **Companions:** [Interpreting a factor-zoo paper →](/guides/interpret-factor-zoo-paper/) · [What makes a factor tradable? →](/guides/what-makes-a-factor-tradable/) · [The backtest checklist →](/guides/evaluate-trading-backtest/) ---------------------------------------------------------------------- # How to Evaluate a Machine-Learning Trading Paper URL: https://thequant.space/guides/evaluate-ml-trading-paper/ Section: Guides Date: 2026-09-07 Description: A checklist for evaluating ML trading papers: baseline honesty, leakage-prone pipelines, accuracy-vs-P&L confusion, seed variance, and the questions that separate signal from citation bait. **Definition.** An ML trading paper claims a learned model — trees, LSTMs, transformers, anything with a loss function — predicts returns, prices, or market states well enough to matter. It is the [largest genre in our archive](/topics/machine-learning/) and the one where the gap between reported and real performance is widest, because ML adds new failure channels *on top of* the [standard backtest sins](/guides/evaluate-trading-backtest/). ## The worked example: the accuracy illusion A paper reports 57% directional accuracy on daily S&P moves. Sounds tradeable. Now the arithmetic: with daily vol ~1%, a 57/43 edge on direction — *if real* — yields a gross Sharpe near 2, but three deductions apply before belief. (1) The base rate: the index rises ~53% of days, so a constant "up" prediction already scores 53% — the paper's edge over *naive* is 4 points, not 7. (2) Accuracy weights all days equally; P&L doesn't — models routinely get the small moves right and the crashes wrong, so accuracy can rise while returns fall. (3) The number is one seed, one split. Rerun with five seeds and the spread often covers the entire claimed edge. Any paper reporting accuracy without a naive-baseline comparison, a P&L translation, *and* seed variance has reported approximately nothing. ## The ML-paper checklist - ☐ **Baselines that can win**: logistic regression, gradient-boosted trees on the same features, and buy-and-hold. A deep model beating only other deep models is an ablation. Trees-on-tabular remains the bar to clear in this field — a result that [recurs across our hub](/topics/machine-learning/). - ☐ **Pipeline-level leakage audit**: scalers fit on the full sample, features selected by full-sample correlation, labels overlapping the [train/test boundary](/guides/walk-forward-out-of-sample-testing/#why-the-ml-playbook-fails-on-market-data), pretrained models with [post-test-period knowledge](/guides/look-ahead-bias-point-in-time-data/#7-llm-training-data-leakage--the-newest-member). ML leakage lives in the pipeline, not the dataset. - ☐ **Temporal splits with purging/embargo** — random k-fold on market data is disqualifying, full stop. - ☐ **Seeds and stability**: multiple seeds with dispersion reported; single-run deep-learning results in finance are anecdotes. - ☐ **Hyperparameter accounting**: the search grid is a [trial count](/guides/deflated-sharpe-ratio/); tuned-on-test is the genre's quiet epidemic. - ☐ **Economic translation**: predictions → portfolio → [costs](/guides/transaction-costs-slippage-market-impact/) → net Sharpe. Papers stopping at RMSE or accuracy have tested a model, not a strategy. - ☐ **Was the baseline trained correctly?** Verify, don't assume — [our T-KAN replication review](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion) found a headline comparison whose baseline LSTM was never trained at all. ## Questions to ask while reading 1. What does the naive baseline score, and is the gap economically meaningful after costs? 2. Could I reproduce the *data pipeline* from the paper alone — or only the architecture? 3. How many configurations were evaluated, and where did the test set enter that loop? 4. Does performance concentrate in a subperiod, a volatility regime, or a handful of names? 5. If the model is interpretable-by-claim, do the shown "learned" artifacts actually change with training? (We've caught fixed activation functions labeled as learned.) ## What our scoring captures — and misses ML papers cluster in our "[Holy Grail](/score-guide/)" and "Lab Rats" quadrants: the rigor axis rewards real out-of-sample protocol and cost accounting, and it demotes the accuracy-only genre effectively. What the score can't verify from a paper's text is *pipeline* leakage — a paper can describe a clean protocol and implement a dirty one. That's only catchable at the code level, which is why ML papers are the genre where [reproduction](/guides/reproduce-quant-research/) pays the highest information dividend per hour spent. --- **Companions:** [Evaluating RL trading papers →](/guides/evaluate-rl-trading-paper/) · [Walk-forward & OOS →](/guides/walk-forward-out-of-sample-testing/) · [Backtest overfitting →](/guides/backtest-overfitting/) ---------------------------------------------------------------------- # How to Evaluate a Market-Microstructure Paper URL: https://thequant.space/guides/evaluate-microstructure-paper/ Section: Guides Date: 2026-09-07 Description: A checklist for evaluating market-microstructure papers: dataset provenance, venue and period specificity, theoretical assumption audits, and the generalization trap. **Definition.** A [market-microstructure paper](/topics/market-microstructure/) studies how prices form: limit order books, market making, impact, adverse selection, liquidity. The genre splits cleanly into **theory** (stochastic models of books and dealers) and **empirics** (measurement on tick data), and each half fails differently — so the first evaluation step is deciding which paper you're holding, because the checklists barely overlap. ## The worked example: true here, false there An empirical paper measures order-book-imbalance predictability: top-of-book imbalance predicts the next mid-move with 65% accuracy on NASDAQ large caps, 2014, ITCH data. Every word of that sentence is a scope condition. The same measurement on small caps (wider ticks, sparser books), on 2024 data (different HFT ecology), on a crypto CLOB (different fee/queue mechanics), or from a consolidated SIP feed (different timestamps, no queue detail) can shrink toward coin-flip. The paper is *correct*; the reader who deploys it on Binance perpetuals is not. Microstructure results are the most venue- and era-local in all of quant finance — generalization is the reader's risk, not the author's claim. ## The empirical checklist - ☐ **Dataset provenance first**: exchange-native feed with order IDs (ITCH/OUCH-class), consolidated tape, or vendor-processed book snapshots? Results about queues and cancellations require the first; papers inferring them from trades-and-quotes are estimating, not measuring. - ☐ **Venue, universe, period stated and *narrow***: treat every result as tagged with them. Check whether the sample predates structural breaks (maker-taker changes, tick-size regimes, the venue's HFT maturity). - ☐ **Timestamp discipline**: exchange vs receipt time, clock sync across feeds — at these horizons [timestamp semantics](/guides/look-ahead-bias-point-in-time-data/#5-timestamp-semantics) *is* the methodology. - ☐ **Microstructure noise handled**: bid-ask bounce and discreteness masquerade as predictability and inflate high-frequency [volatility estimates](/topics/volatility/); look for noise-robust estimators or mid-quote (not trade-price) construction. - ☐ **Economic translation with maker/taker realism**: a 55% next-tick edge is worthless crossing the spread and possibly viable providing liquidity — which flips the relevant [cost model](/guides/transaction-costs-slippage-market-impact/) from taker fees to queue position and adverse selection. ## The theory checklist - ☐ **Assumption audit as the whole game**: Poisson arrivals, exponential fills, constant spreads, no feedback from the agent's own quotes — each is a known counterfactual. The question isn't purity; it's *which assumption, when relaxed, kills the conclusion*. - ☐ **Calibration section present?** Theory papers that fit their model to any real book earn a rigor tier above those that don't; "Lab Rats" [quadrant placement](/score-guide/) is the default otherwise. - ☐ **Testable prediction extracted**: the best theory papers imply a measurement someone could run. If you can't state one, the paper is mathematics with market-flavored notation — sometimes valuable, never evidence. ## Questions to ask while reading 1. Would this result survive on a different venue, era, or fee structure — and does the paper test even one? 2. What data would I need to reproduce this, and [can it be bought](/guides/market-data-vendors/)? 3. For prediction claims: is the horizon before or after the venue's round-trip latency makes it actionable? 4. For theory: which single assumption does the conclusion lean on hardest? 5. Is "liquidity" defined as spread, depth, resiliency, or impact — and consistently? ## What our scoring captures — and misses The [quadrant system](/score-guide/) was practically designed for this genre: the theory half lands in Lab Rats, the measurement half in Street Traders, and papers doing both — model plus calibration — surface as Holy Grail. The structural blind spot is **scope**: rigor scores reward how well a paper establishes its claim, not how far the claim travels, and microstructure claims travel worse than any other genre's. Read the hub with venue-and-period tags in your head; the score tells you the paper is sound, not that it's *yours*. --- **Companions:** [LOB imbalance: measures and misses →](/guides/limit-order-book-imbalance/) · [Market making mechanics →](/guides/market-making-mechanics/) · [HFT & execution hub →](/topics/high-frequency-trading/) ---------------------------------------------------------------------- # How to Evaluate a Reinforcement-Learning Trading Paper URL: https://thequant.space/guides/evaluate-rl-trading-paper/ Section: Guides Date: 2026-09-07 Description: A checklist for evaluating RL trading papers: environment leakage, reward hacking, the sample-efficiency problem in non-stationary markets, and where RL claims are actually credible. **Definition.** An RL trading paper trains an agent to act — allocate, execute, quote, hedge — by interacting with a market environment and maximizing cumulative reward. The framing is seductive because it skips the forecast-then-optimize pipeline; the evaluation problem is that markets are close to the hardest possible RL setting — non-stationary, near-random-walk, weak reward signal — so [our RL hub](/topics/reinforcement-learning/) contains both genuine advances and the field's most spectacular simulator victories. ## The worked example: winning a game nobody plays A DQN portfolio agent reports 3× buy-and-hold on five years of daily crypto data. Deconstruct the environment: the agent observes today's OHLCV *including the close*, acts "at" that close ([same-bar execution](/guides/look-ahead-bias-point-in-time-data/#4-same-bar-execution)), pays zero spread in an asset class where [taker costs are 2–10 bps+](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges), and its reward includes unrealized mark-to-market on positions sized without impact. Each simplification is defensible in an RL-methods paper; jointly they define a game whose optimal policy has no market counterpart. The agent won the game. The game was the leak. ## The RL-paper checklist Everything on the [ML checklist](/guides/evaluate-ml-trading-paper/) applies. Additionally: - ☐ **Environment audit before results**: observation timing vs action timing, cost and impact model, and whether rewards use information the agent couldn't have. In RL papers *the environment is the methodology* — read it the way you'd read a backtest engine. - ☐ **Dumb-baseline comparison**: buy-and-hold, equal-weight, TWAP (for execution). RL papers benchmarking only against other RL agents are comparing losers of different games. - ☐ **Train/test regime separation**: an agent trained through 2020–21 crypto and tested on 2022 tells you something; trained and tested inside one bull regime tells you nothing. [Regime dependence](/guides/regime-dependence/) is fatal here because policies encode regime implicitly. - ☐ **Seeds, and more than ML needs**: RL variance across seeds is notoriously large; fewer than 5 seeds with dispersion is a screenshot. - ☐ **Sample-efficiency honesty**: how many environment steps did training take, and how many *independent* market periods does that actually represent? A million steps over one replayed year is one year, replayed. - ☐ **Reward-hacking check**: does the reward function admit degenerate optima (e.g., volatility-pumping unrealized P&L, or churning for a per-trade bonus)? ## Where RL claims deserve real attention Calibrate skepticism by sub-domain. **Execution and market making** — dense feedback, clear cost structure, stationary-enough microstructure — is where RL results replicate and deploy; the [optimal-execution literature](/topics/high-frequency-trading/) meets RL credibly. **Hedging** (deep hedging under transaction costs) is the second defensible island. **End-to-end portfolio alpha from prices** is where the sample-efficiency arithmetic is unforgiving; hold those papers to the maximum standard. ## Questions to ask while reading 1. Could the environment's optimal policy be profitable in a real market, even in principle? 2. What fraction of the reported gain survives next-bar execution and realistic costs? 3. How many truly independent market regimes did the agent train across? 4. Is the comparison seed-matched, cost-matched, and [correctly trained](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion) on the baseline side? 5. Does the paper report *policy behavior* (turnover, exposure, inventory) or only returns? Behavior is where hacking shows. ## What our scoring captures — and misses The [rigor axis](/score-guide/) treats environment realism as empirical rigor, so RL papers with honest cost/impact modelling and multi-seed reporting surface at the top of the hub. The blind spot is symmetric with ML's: a described environment can differ from the implemented one, and reward hacking is often only visible in training curves and logged behavior a paper omits. When an RL result matters to you, the environment code *is* the paper — [reproduce it](/guides/reproduce-quant-research/) or don't rely on it. --- **Companions:** [ML trading papers →](/guides/evaluate-ml-trading-paper/) · [Regime dependence →](/guides/regime-dependence/) · [Backtest checklist →](/guides/evaluate-trading-backtest/) ---------------------------------------------------------------------- # How to Find Code for Finance Research Papers URL: https://thequant.space/guides/find-code-finance-papers/ Section: Guides Date: 2026-09-07 Description: Locating implementations of quant finance papers: official repos, third-party reimplementations and their risks, the library ecosystem, and how to audit found code before trusting it. **The problem.** Finance has no strong code-publication norm: most papers ship no implementation, official repos often produce "illustrative" results rather than the paper's tables, and the third-party reimplementations that fill the gap [inherit and add bugs](/guides/reproduce-quant-research/#stage-3--reimplement-independently-on-purpose). Finding code is easy; finding code you should *trust* is the actual skill. ## The four tiers, and where to look **Tier 1 — Official code.** Check, in order: the paper's abstract/footnotes for a repo link; Papers With Code (thin for finance, improving for [ML-finance crossovers](/guides/map-ml-asset-pricing/)); authors' GitHub profiles and academic homepages (code sometimes posts *after* publication — check again in six months); and journal supplementary materials, where "replication package" requirements at some finance journals quietly deposit real code. Pin the commit hash the moment you find it. **Tier 2 — Third-party reimplementations.** GitHub-search the paper title in quotes, the method's distinctive name ("DeepLOB", "Avellaneda-Stoikov", "deflated Sharpe"), or its equation's distinctive term. These are simultaneously the ecosystem's gift and its trap: an [audit of one such notebook](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion) — untrained baseline, label-as-return backtest, inverted labels — is this site's standing exhibit on why stars ≠ correctness. Popular reimplementations of famous papers have propagated known bugs for *years*. **Tier 3 — Library implementations.** The method may live, hardened, in the ecosystem already: statsmodels (econometrics), arch (GARCH family), [QuantLib](/resources/) (pricing), skfolio/PyPortfolioOpt (allocation), mlfinlab-descendants (López de Prado methods — [CSCV](/guides/cscv-explained/), purged CV), vectorbt (backtesting). A maintained library beats a dissertation repo on every axis except exact-paper fidelity — and fidelity is checkable. **Tier 4 — Nothing exists.** For most finance papers, this is the tier. Then the paper's [flowchart-level pipeline](/flowcharts/) is your spec, and writing the implementation *is* the reproduction — with the [known-answer tests](/guides/auditing-a-monte-carlo/#61-test-invariants-not-outputs) that entails. ## The 15-minute audit before found code touches your research - ☐ **Does it run and reproduce a paper number?** Not "runs without error" — produces a *table entry* within tolerance. Illustrative repos fail this in the first ten minutes and save you the next ten hours. - ☐ **Grep the red flags**: seeds unset, `strict=False`-style silent loading, [full-sample fitting before splits](/guides/look-ahead-bias-point-in-time-data/#6-normalization-and-statistics-computed-on-the-full-sample), test data touched during training, hardcoded lookahead indices. - ☐ **Check the environment**: no lockfile/pinned versions means results drift with dependency updates — pin before running, per [reproduction stage 3](/guides/reproduce-quant-research/#stage-3--reimplement-independently-on-purpose). - ☐ **Read the issues tab**: unanswered "results don't match the paper" issues are the most informative documents on GitHub. - ☐ **Diff against the paper's methodology**, not its README. Reimplementations drift toward what was easy, and READMEs describe intentions. ## Using found code without inheriting its conclusions The disciplined pattern: use third-party code to **understand** (reading a working implementation clarifies ambiguous notation faster than the paper), but validate independently before any result enters your [experiment log](/guides/quant-research-pipeline/) — known-answer tests, [one-bar-lag checks](/guides/evaluate-trading-backtest/#2-could-the-signal-have-known-the-future), and reconciliation against the paper's own numbers. Code you didn't audit is an assumption you didn't list. ## Questions to ask of any repo 1. Which paper table can this code reproduce, and has anyone (issues, forks) confirmed it? 2. What's the commit history — active refinement, or a graduation-day snapshot? 3. Where does the data come from, and [can you get it](/guides/find-datasets-quant-papers/)? 4. If it's a library: does its implementation match the paper's variant, or a simplified cousin? (Check the docstring's citations.) ## How our scoring treats code availability Papers with linked, working code score higher on [rigor's reproducibility component](/score-guide/) — code is the strongest single falsifiability signal a paper can send. The score can't audit the code itself (that's tier-2 caution and yours to apply), which is why paper pages carry [discussion sections](/guides/reproduce-quant-research/#stage-6--document-make-your-reproduction-reproducible): replication notes on real repos are precisely the knowledge the field undersupplies. --- **The chain:** [find the data →](/guides/find-datasets-quant-papers/) · [reproduce properly →](/guides/reproduce-quant-research/) · [what to replicate first →](/guides/choose-papers-to-replicate/) ## Papers in the archive that link code The shortlist below is generated from the archive: every paper whose page links a public repository, ranked by rigor-weighted score. The full, daily-updated list lives at [papers with code](/papers-with-code/), and the **Has code** filter in [search](/search/) combines it with topic and rigor filters. ---------------------------------------------------------------------- # How to Find the Datasets Used in Quant-Finance Papers URL: https://thequant.space/guides/find-datasets-quant-papers/ Section: Guides Date: 2026-09-07 Description: Locating the data behind quant finance papers: the standard dataset zoo (CRSP, TAQ, FI-2010, Kaggle mirrors), decoding data sections, access tiers, and legitimate substitutes when the original is unreachable. **The problem.** [Reproduction](/guides/reproduce-quant-research/) dies at the data step more often than at any other — not because data doesn't exist, but because papers describe it vaguely ("daily US equity data"), access it institutionally (WRDS subscriptions), or used it under licenses you can't inherit. This guide is the location layer: what the standard datasets are, where they live, and what substitutes are legitimate when the original is out of reach. ## The standard-dataset zoo A surprising share of the literature runs on a small set of recurring datasets — recognize them on sight: - **CRSP / Compustat** (via WRDS): the substrate of US [asset-pricing research](/topics/factor-investing/) — survivorship-clean equities and point-in-time-ish fundamentals. Access: academic affiliation (cheap through a university library) or institutional money. No free substitute matches its [delisting coverage](/guides/survivorship-bias/#building-a-survivorship-clean-backtest). - **TAQ**: US trades and quotes, the [microstructure](/topics/market-microstructure/) workhorse. Institutional. Modern papers increasingly use exchange-native ITCH data instead — even less accessible, more precise. - **FI-2010**: the benchmark [LOB dataset](/guides/map-lob-prediction/) — free, downloadable, and the reason dozens of deep-LOB papers are comparable at all. Know its quirks: z-scored prices (mid-price reconstruction is painful — [a problem we hit ourselves](/guides/auditing-a-monte-carlo/)), one week, ten Finnish stocks. Results on FI-2010 are results *about FI-2010*. - **Ken French's data library**: factor returns and sorted portfolios, free, the standard benchmark layer for [factor papers](/guides/evaluate-factor-investing-paper/). - **Option surfaces (OptionMetrics)**: the [options-research](/topics/options-derivatives/) standard; institutional. Free substitutes are thin — this is the hardest genre to reproduce cheaply. - **Crypto**: the happy exception — exchange APIs give full public history free, which is why [crypto papers](/topics/crypto-defi/) are the most reproducible genre in the archive. ## Decoding a vague data section When the paper doesn't name its source, triangulate: the **citation trail** (data citations hide in footnotes and acknowledgments — "we thank X for providing…" is a data source); the **summary-statistics fingerprint** (universe size per year, date ranges, and mean returns narrow candidates fast — [matching these is stage 4 of reproduction anyway](/guides/reproduce-quant-research/#stage-4--reconcile-match-statistics-before-results)); the **appendix of a *later* paper** (replications and follow-ups often document the original's data more precisely than the original did); and the **authors** — a short, specific email asking "which vendor and which adjustment settings?" gets answered more often than cynics expect, especially within a year of publication. ## The access-tier reality | Tier | What it unlocks | Cost reality | |---|---|---| | Free (arXiv-adjacent) | FI-2010, French library, crypto APIs, FRED, some Kaggle mirrors | $0 — but mirrors are unversioned; verify against summary stats | | Individual practitioner | Daily/intraday prices via [API vendors](/guides/market-data-vendors/) | $30–500/mo — covers price-based reproduction | | Academic | WRDS: CRSP, Compustat, TAQ, OptionMetrics | University affiliation; the cheapest legitimate path to research-grade history | | Institutional | Exchange-native feeds, PIT fundamentals at scale | Four figures monthly and up | The practical implication for independents: **price-based papers are reproducible; fundamentals- and options-based papers usually require the academic tier**. Factor that into [which papers you choose to replicate](/guides/choose-papers-to-replicate/). ## Substitution rules — when exact data is unreachable Substituting data changes the experiment; do it honestly: - ☐ Match the *properties that drive the result*, not the label: frequency, universe breadth, [survivorship handling](/guides/survivorship-bias/), adjustment policy. - ☐ Reconcile summary statistics between substitute and original before comparing results — divergence here explains divergence everywhere. - ☐ Reframe your conclusion: "the effect replicates on Polygon data 2015–2025" is a *generalization test*, not a reproduction — arguably [more valuable](/guides/robustness-trading-strategy/#2-universe-robustness--the-transfer-test), but a different claim. - ☐ Never substitute silently in a write-up; the substitution *is* a finding. ## Questions to ask of any paper's data 1. Could I buy this data today, at which [tier](/guides/market-data-vendors/)? 2. Is the dataset one of the standard zoo (comparable results exist) or bespoke (nothing to benchmark against)? 3. What does the choice of dataset *conveniently exclude* — [dead names](/guides/survivorship-bias/), a crisis, an asset class? 4. If it's a public benchmark like FI-2010: is the paper state-of-the-art on the benchmark, or on the problem? ## How our scoring reflects data accessibility Data transparency feeds the [rigor score](/score-guide/) — named sources, date ranges, and construction rules score above "proprietary dataset" every time, and the paper pages surface each paper's data description in the abstract and flowchart. What the score deliberately doesn't do is penalize institutional data (TAQ-based work can be superb); it penalizes *vagueness*. Accessibility is your constraint to apply — via the tier table above. --- **The chain continues:** [find the code →](/guides/find-code-finance-papers/) · [reproduce it →](/guides/reproduce-quant-research/) · **Data infrastructure:** [vendor guide →](/guides/market-data-vendors/) ---------------------------------------------------------------------- # How to Interpret a Factor-Zoo Paper URL: https://thequant.space/guides/interpret-factor-zoo-paper/ Section: Guides Date: 2026-09-07 Description: Reading factor-zoo and replication meta-studies: what the multiple-testing corrections mean, why replication rates differ so wildly between studies, and what survives for practitioners. **Definition.** A factor-zoo paper is a meta-study: instead of proposing an anomaly, it audits hundreds of published ones under a single methodology — common data, common construction, and [multiple-testing corrections](/guides/statistical-vs-economic-significance/#why-the-usual-significance-bar-is-broken-in-finance) sized to the *field's* collective search. The genre's landmarks — Harvey-Liu-Zhu's t ≈ 3 hurdle, Hou-Xue-Zhang's mass replication, McLean-Pontiff's post-publication decay estimates — are among the most consequential papers in [asset pricing](/topics/factor-investing/), and among the most misread. ## Why honest studies disagree so much Replication rates across factor-zoo studies range from grim (~35–50%) to reassuring (~80%+) *on overlapping factor sets*. The disagreement is methodological, and decoding it is the whole skill: - **Construction choices**: [equal- vs value-weighting, micro-cap inclusion, breakpoints](/guides/evaluate-factor-investing-paper/#the-worked-example-construction-is-the-hidden-factor) — the harsh studies use constructions that mute micro-cap noise; the generous ones follow original papers' choices. Both are defensible; they answer different questions ("is there a tradeable premium?" vs "did the authors report their sample accurately?"). - **The correction's severity**: a Bonferroni-style correction across 400 factors is brutal; false-discovery-rate control is milder; hierarchical/Bayesian approaches milder still. The reported replication rate is largely a function of this dial. - **The universe of factors counted**: auditing every published anomaly vs the subset with strong priors changes denominators wholesale. - **Sample updates**: extending samples past publication mixes replication failure with [genuine decay](/guides/regime-dependence/#the-live-trading-diagnosis-problem) — different diagnoses, same symptom. Reading rule: **extract the study's answer to "replicates under whose construction, corrected how?" before quoting its headline rate.** A factor "failing" in one study and "passing" in another is usually both — under different, stated rules. ## What survives everyone's methodology The convergent findings across the genre, stated conservatively: a **small core** of factors (market, profitability/quality-family, momentum-family, value-family in some constructions) clears even harsh corrections; **post-publication returns decay substantially** — McLean-Pontiff-style estimates put the decline around a third to a half, blending [arbitrage and reversal of data-mining luck](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer); **micro-caps carry the zoo** — a large share of anomalies live where [costs](/guides/transaction-costs-slippage-market-impact/) and [survivorship issues](/guides/survivorship-bias/) concentrate and capacity doesn't; and **correlated factors are few underlying bets** — hundreds of published factors compress into a handful of statistical clusters, so "N factors replicate" overstates independent evidence exactly the way [correlated trades overstate sample size](/guides/evaluate-trading-backtest/#4-how-many-independent-bets-is-this-really). ## The practitioner translation - A zoo paper is the *best prior generator* available: before evaluating any [new factor paper](/guides/evaluate-factor-investing-paper/), know which cluster it belongs to and how that cluster fared under harsh construction. - Treat "fails harsh replication" as *capacity information*, not falsity: an equal-weighted-only anomaly is a statement about where the effect lives. - The decay estimates are your default haircut: whatever premium you estimate from a published factor's history, the zoo literature says to expect a large fraction less going forward. ## The checklist for reading one - ☐ Construction rules and weighting stated and understood - ☐ Correction method identified; severity dial noted - ☐ Replication rate re-read as conditional on both - ☐ Factor-cluster structure reported (or reconstruct it) - ☐ In-sample vs post-publication failure distinguished - ☐ The study's own [trial count considered](/guides/deflated-sharpe-ratio/) — meta-studies make choices too ## Questions to ask 1. Under this study's rules, would the *market factor itself* replicate? (A useful severity calibration.) 2. Which specific factors flip status relative to other zoo studies, and does the construction difference explain it? 3. Does the study distinguish "never real" from "real, then arbitraged"? 4. What does it imply for the factor *cluster* behind the paper you actually care about? ## How our scoring relates Zoo papers themselves tend to score at the top of our [rigor axis](/score-guide/) — they are, in effect, the field auditing itself with disclosed methodology, the same ethos as [our replication reviews](/guides/reproduce-quant-research/). For the individual factor papers they audit, our scores and the zoo's verdicts are complementary: we score the paper's internal evidence quality; the zoo scores its survival under external, uniform rules. A paper strong on both is the rare real thing. --- **The framework upstream:** [multiple testing & significance →](/guides/statistical-vs-economic-significance/) · [evaluating factor papers →](/guides/evaluate-factor-investing-paper/) · **The mechanism:** [what makes a factor tradable →](/guides/what-makes-a-factor-tradable/) ---------------------------------------------------------------------- # How to Interpret Confidence Intervals in Strategy Research URL: https://thequant.space/guides/confidence-intervals-strategy-research/ Section: Guides Date: 2026-09-07 Description: Reading confidence intervals on Sharpe ratios, alphas, and backtest statistics: what they actually say, how wide they really are, block bootstrapping, and the selection problem that breaks naive intervals. **Definition.** A 95% confidence interval around a strategy statistic says: *under this procedure, intervals built this way cover the true value 95% of the time.* It is a statement about the estimator's noise, not a probability statement about your strategy — and in strategy research its two chronic problems are that honest intervals are **shockingly wide**, and published ones are **systematically too narrow**, because selection happened before the interval was drawn. ## The numbers nobody internalizes The standard error of an annualized Sharpe estimated from T years of (roughly i.i.d.) daily data is approximately **√((1 + SR²/2)/T)** — for practical purposes, near **1/√T** for moderate Sharpes. Turn that into intervals: | Track length | 95% CI on a measured SR of 1.0 | |---|---| | 1 year | −1.0 to 3.0 | | 4 years | 0.0 to 2.0 | | 10 years | ~0.36 to 1.64 | | 25 years | ~0.6 to 1.4 | Read the second row again: after four years, "SR = 1" is statistically indistinguishable from zero at conventional confidence. Consequences: comparing two strategies with |ΔSR| < ~0.5 on a few years of data is [reading noise](/guides/evaluate-ml-trading-paper/#the-worked-example-the-accuracy-illusion); manager evaluation on 3-year windows is astrology with spreadsheets; and any paper drawing conclusions from second-decimal Sharpe differences has confused precision with information. Non-normality widens everything further — negative skew and fat tails make the effective interval [wider than the Gaussian formula admits](/guides/high-sharpe-ratio-not-investable/#3-volatility-isnt-risk--skew-is). ## Computing intervals honestly: the block bootstrap Market returns are autocorrelated and volatility-clustered, so the i.i.d. bootstrap understates uncertainty. The standing tool is the **block bootstrap** (or stationary bootstrap with geometric block lengths — [the generator menu](/guides/auditing-a-monte-carlo/#41-the-generator-menu)): ```python def block_ci(rets, stat, n_boot=2000, block=21): T = len(rets); n_blocks = T // block + 1 out = [] for _ in range(n_boot): starts = rng.integers(0, T - block, n_blocks) sample = np.concatenate([rets[s:s+block] for s in starts])[:T] out.append(stat(sample)) return np.percentile(out, [2.5, 97.5]) ``` Block length should cover the dependence horizon (weeks for daily equity returns; longer for [regime-sensitive strategies](/guides/regime-dependence/)). The same machinery gives intervals on drawdowns, quantiles, and anything else [without a closed form](/guides/auditing-a-monte-carlo/#52-quantiles-have-no-closed-form--bootstrap-them). For simulation outputs, remember the estimate itself has [sampling error worth reporting](/guides/auditing-a-monte-carlo/#51-standard-error-of-a-simulated-probability). ## The selection problem: why published intervals lie A confidence interval is valid for a *pre-specified* statistic. The interval around your **best-of-N** backtest is a different object entirely: the selection already consumed the luck the interval thinks is still available, so naive CIs on selected results are too narrow — often absurdly so. The corrections live in the [deflated Sharpe](/guides/deflated-sharpe-ratio/) / [CSCV](/guides/cscv-explained/) family; the reading rule is simpler: **an interval is only as honest as the trial count behind it**, and papers reporting tight intervals on searched results have reported the wrong distribution. ## The checklist - ☐ Interval computed with dependence-aware methods (block bootstrap / HAC errors), not i.i.d. formulas - ☐ Width sanity-checked against the 1/√T table — suspiciously tight intervals mean a method error - ☐ Statistic pre-specified, or interval adjusted for [selection](/guides/backtest-overfitting/) - ☐ Skew/kurtosis reported alongside (the interval's shape matters, not just width) - ☐ Comparisons between strategies use the interval on the *difference*, with common time periods ## Questions to ask when reading a paper 1. Is any interval or standard error reported at all — and if not, what would the 1/√T arithmetic say about the headline number? 2. What dependence structure did the errors assume, and does the strategy's holding period violate it? 3. Was this statistic selected before or after seeing results? 4. Do the paper's conclusions survive at the interval's unfavorable edge? ## What our scoring captures — and misses Reporting uncertainty at all is rare enough that it moves papers up the [rigor axis](/score-guide/) — intervals, standard errors, and bootstrap sections are among the strongest textual signals of care we score on. The miss is the selection problem: a beautifully computed interval on a silently selected result scores well and misleads anyway. Pair the score with the trial-count question, always. --- **The width arithmetic's sequel:** [Deflated Sharpe →](/guides/deflated-sharpe-ratio/) · **The dependence machinery:** [Auditing a Monte Carlo →](/guides/auditing-a-monte-carlo/) · **The comparison discipline:** [significance guide →](/guides/statistical-vs-economic-significance/) ---------------------------------------------------------------------- # How to Read a Quant Finance Paper (Without Wasting Your Afternoon) URL: https://thequant.space/guides/how-to-read-quant-finance-papers/ Section: Guides Date: 2026-09-07 Description: A triage framework for reading quantitative finance papers: the 10-minute pass, the two axes that matter, section-by-section priorities, and the red flags that end a read early. arXiv's q-fin sections and SSRN publish thousands of papers a year. Reading them linearly — abstract, introduction, literature review, model, results — is how an afternoon disappears into a paper that a 10-minute triage would have rejected. This is the reading order we use to process [5,000+ papers](/flowcharts/), and the framework behind our [scoring system](/score-guide/). ## The two axes that decide everything Before any technique: know what question you're asking. Almost every judgment about a quant paper collapses onto two independent axes: - **Math complexity** — how much theoretical machinery the paper runs on: stochastic calculus and PDEs at one end, descriptive statistics at the other. - **Empirical rigor** — how seriously the claims were tested: real data, out-of-sample discipline, transaction costs, robustness checks. These are *independent*. Elegant theory with zero data ("Lab Rats" in our [quadrant system](/score-guide/)) is valuable for pricing engines and dangerous for trading claims. A simple idea tested brutally ("Street Traders") is often more tradeable than a beautiful one tested never. The most common reading error is letting the math axis stand in for the rigor axis — difficulty is not evidence. ## The 10-minute pass, in the right order Read sections in order of *information per minute*, not page order: **1. Abstract (1 min):** extract exactly two things — the claim, and whether it's a theory claim or a performance claim. Performance claims ("Sharpe of 2.1") set up the rest of the triage; theory claims redirect you to a different standard entirely. **2. The data section (3 min):** this is the real abstract. Which instruments, which dates, which source, which frequency? Three instant filters: Does the sample end years ago (why?)? Is the universe survivorship-clean ([the bias that manufactures alpha](/guides/survivorship-bias/))? Is the data purchasable by anyone, or proprietary — reproducibility starts here. **3. The main results table (3 min):** go straight to it. Look for what's *missing* before what's present: transaction costs, turnover, out-of-sample columns, a max drawdown. A returns table without turnover is unpriceable — see [our cost framework](/guides/transaction-costs-slippage-market-impact/). Check the number of observations against the number of things tested (the [multiple-testing problem](/guides/statistical-vs-economic-significance/)). **4. The methodology's one load-bearing paragraph (2 min):** every paper has one — usually where the signal is defined or the estimator chosen. Find it and ask the question from our [Monte Carlo audit](/guides/auditing-a-monte-carlo/): *is each parameter assumed or measured?* **5. Skip entirely on the first pass:** the literature review, the robustness appendix (note that it exists — its absence is data), and the conclusion, which restates the abstract with more confidence than the results earned. Ten minutes in, you can place the paper on both axes and decide: read deeply, mine for one idea, or close. ## Red flags that end a read early - **Backtest starts when the data conveniently starts** — sample-period-as-modelling-assumption, never acknowledged. - **Performance claims with no costs** and no reply to the obvious objection. Fixable by you, but the omission signals what else was skipped — run it through [the backtest checklist](/guides/evaluate-trading-backtest/). - **Accuracy above 60% on daily direction** of liquid instruments — almost always [look-ahead leakage](/guides/look-ahead-bias-point-in-time-data/) or an illiquid universe. - **One spectacular equity curve, no error bars, no subperiods.** A single path is one draw from a fan — [see it interactively](/tools/equity-curve-simulator/). - **"Novel" + "AI" + no baseline.** A deep model beating no sensible benchmark is an ablation, not a result. ## Red flags that *aren't* Symmetry demands the opposite list. Don't discard a paper for: modest Sharpe ratios (0.7 honestly measured beats 2.5 with leakage); negative or null results (the rarest and most trustworthy findings in the field); simple methods (simplicity survives out of sample better than complexity — a finding that recurs across our [machine-learning hub](/topics/machine-learning/)); or prose quality (some of the best empirical work reads terribly). ## Where the deep read goes, when a paper earns it Deep reading is reverse engineering, not comprehension: reconstruct the pipeline — data → transformation → signal → portfolio → evaluation — until you could reimplement it. That's precisely what our [flowcharts](/flowcharts/) draw for each paper, and why they exist: the pipeline diagram is the paper's skeleton, and most flaws are visible in the skeleton. If the paper matters enough to trade on, the next step is the [reproduction checklist](/guides/reproduce-quant-research/) — because the read-through is the cheapest gate, not the last one. --- **Browse triaged papers:** [by topic](/topics/) · [highest empirical rigor first](/) · **The scoring rubric:** [score guide →](/score-guide/) ---------------------------------------------------------------------- # How to Search Quantitative-Finance Literature Efficiently URL: https://thequant.space/guides/search-quant-finance-literature/ Section: Guides Date: 2026-09-07 Description: A working system for searching quant finance literature: where each source excels (arXiv, SSRN, Google Scholar), citation-chain tactics, freshness triage, and search operators that actually work. **Definition of the problem.** Quant finance literature is split across venues with different cultures: **arXiv q-fin** (fast, math/ML-heavy, uneven quality), **SSRN** (empirical finance and asset pricing, often years before journals), and **journals** (slow, peer-reviewed, paywalled — and the version of record that sometimes differs from the preprint you read). Efficient search means using each for what it's good at and letting citation structure, not keywords, do the heavy lifting. ## The source map | Need | Best source | Tactic | |---|---|---| | What's new this week | arXiv q-fin new/recent listings | Or let [our daily pipeline](/posts/) triage it — that's literally what it's for | | The empirical asset-pricing canon | SSRN + Google Scholar | Sort SSRN by downloads within a topic; downloads ≈ practitioner attention | | Who settled this question | Google Scholar cited-by chains | See citation-chaining below | | Implementable methods | arXiv cs.LG/stat.ML crossovers | Search the method name + "trading"/"portfolio" | | The published version of record | Journal sites / author pages | Diff the tables against the preprint before citing numbers | ## Citation-chaining beats keyword search Keywords find papers; citations find *literatures*. The method: locate one strong anchor paper (a survey, or a high-[rigor](/score-guide/) empirical piece), then walk **backward** through its references for foundations and **forward** through Google Scholar's "cited by" for everything that built on it — filtering the forward pass by recency and citation count. Two hops from a good anchor covers a subfield more completely than any keyword session, and it surfaces the *critiques* (papers citing the anchor to disagree), which keyword search structurally misses and which are [the most informative documents in any literature](/guides/p-hacking-financial-research/). Finding the anchor is what our [research maps](/guides/map-ml-asset-pricing/) and [topic hubs](/topics/) are built for: each hub ranks its papers by rigor-weighted score, so the top of a hub is a pre-vetted anchor list. ## Search tactics that pay for themselves - **Search the method, not the goal**: "temporal convolutional network limit order book" beats "predict stock prices with AI" by an order of magnitude in signal-to-noise. - **Search the dataset**: papers using the same data ("FI-2010", "TAQ", "CRSP delisting") form natural comparison sets — and shared datasets mean comparable results, [a rarity worth exploiting](/guides/find-datasets-quant-papers/). - **Search the critique**: append "replication", "revisited", "does not", "fails" to any famous result's name. The follow-up literature is where the truth lives. - **Author-follow beats topic-follow**: a handful of consistently rigorous authors per subfield is a better feed than any keyword alert. Build the list from hub-top papers. - **Version-check arXiv papers**: v1 vs v4 table diffs reveal what reviewers forced — [informative in both directions](/guides/reproduce-quant-research/#stage-1--acquire-establish-what-exists). ## The two-hour sweep protocol 1. (10 min) Find the anchor via [maps](/guides/), hubs, or a survey. 2. (30 min) Backward pass: skim the anchor's references; keep ~10 foundations. 3. (40 min) Forward pass: cited-by, sorted by relevance then recency; keep ~10 descendants, deliberately including 2–3 critiques. 4. (30 min) [Ten-minute triage](/guides/how-to-read-quant-finance-papers/#the-10-minute-pass-in-the-right-order) the keepers; most fail. 5. (10 min) Log survivors with one-line claims in your [experiment log](/guides/quant-research-pipeline/#1-idea-intake--the-experiment-log). Done — resist the tab explosion. ## Questions that keep a search honest 1. Am I searching to *learn the answer* or to *find support*? The second is [HARKing's](/guides/p-hacking-financial-research/) research-consumption twin. 2. Have I found the paper that *disagrees* yet? A literature review without a critique is incomplete by construction. 3. Is the newest paper actually better, or just newer? Recency bias in method selection is how last year's architecture becomes this year's [unexamined baseline](/guides/evaluate-ml-trading-paper/). ## How this site fits the workflow The archive is a pre-triaged layer for exactly this process: [search 5,000+ scored papers](/search/) with rigor/quadrant filters, use hub rankings as anchor lists, follow the [weekly digest](/posts/) instead of raw arXiv listings, and check the [maps](/guides/map-statistical-arbitrage/) for field structure before diving. The scores don't replace your judgment — they replace the first hour of it. --- **Next steps in the chain:** [read the survivors →](/guides/how-to-read-quant-finance-papers/) · [find their data →](/guides/find-datasets-quant-papers/) · [choose what to replicate →](/guides/choose-papers-to-replicate/) ---------------------------------------------------------------------- # How to Store Tick Data Efficiently URL: https://thequant.space/guides/store-tick-data-efficiently/ Section: Guides Date: 2026-09-07 Description: Practical tick-data storage: partitioning schemes, columnar formats, compression choices, type discipline, and the layout decisions that make years of ticks queryable on one machine. Tick data's volume problem is real — [rungs 4–6 of the data ladder](/guides/market-data-types/#the-ladder) run gigabytes per instrument-year — but it's a *layout* problem, not a hardware problem: the same data stored well vs badly differs by 5–20× in size and 100× in scan speed. This is the practical layer under the [database comparison](/guides/tick-data-databases/); most single-researcher setups need only this guide's file-based approach. ## The four decisions that matter **1. Partitioning: date first, then maybe symbol.** The universal access pattern is time-ranged; partition Parquet by `date=YYYY-MM-DD/` so every backtest scan touches only its window. Add `symbol=` partitioning *only* if single-name queries dominate — thousands of small files per day (the many-symbols × daily-partition explosion) is the classic self-inflicted wound; for broad-universe daily work, one file per date with a symbol column beats a directory zoo. Rule of thumb: target individual files of 100MB–1GB; below ~10MB, partitioning overhead dominates. **2. Format: columnar, standard, boring.** Parquet remains the [vendor-neutral answer](/guides/quant-research-stack/#layer-2-storage--parquet--duckdb-until-it-hurts): columnar (scans read only needed fields), self-describing, readable by every engine ([DuckDB](/guides/tick-data-databases/), Polars, Spark, ClickHouse ingest). Keep the raw vendor payloads too — compressed, append-only, in the [raw zone](/guides/quant-research-pipeline/#2-data--one-clean-layer-biases-handled-once) — because normalization decisions are revisable only if the pre-normalized truth survives. **3. Compression: zstd, and let sorting do the real work.** Zstd at moderate levels is the current default (better ratios than snappy at similar speed). The bigger lever is **sort order before writing**: ticks sorted by (symbol, timestamp) compress dramatically better than arrival-ordered ones, because deltas shrink — prices become runs of near-identical values, timestamps become small increments. Same data, same codec, 2–5× size difference from sort order alone. **4. Types: the discipline that compounds.** Timestamps as int64 nanoseconds UTC ([never strings, never local time](/guides/look-ahead-bias-point-in-time-data/#5-timestamp-semantics)); prices as int64 in ticks/pips or fixed-point decimal — **not float64** (floats waste bits on impossible values, break exact equality on price levels, and accumulate rounding in [P&L reconciliation](/guides/notebook-to-production/#2-postgres-is-the-source-of-truth-not-your-process-memory)); sizes as integers; categorical columns (symbol, venue, side) dictionary-encoded — automatic in Parquet if the types are right. Schema-drift protection: version your schema, and validate on ingest — [the feed will change under you](/guides/free-vs-paid-market-data/#common-ways-free-data-research-fails-late). ## The reference layout ```text data/ ├── raw/ # vendor payloads, append-only, zstd │ └── source=polygon/date=2026-09-05/*.jsonl.zst ├── ticks/ # normalized, sorted, typed │ └── date=2026-09-05/trades.parquet (symbol, ts, px, sz, venue, cond) ├── quotes/ │ └── date=2026-09-05/nbbo.parquet └── bars/ # derived, regenerable └── freq=1m/date=2026-09-05/bars.parquet ``` Derived layers (bars, features) are *regenerable* — mark them as such and never let irreplaceable and rebuildable data share a backup policy. [Version the transformations](/guides/versioning-datasets-backtests/), not just the outputs. ## Sizing reality check Order-of-magnitude for a liquid US equity, per year: trades ~50–200MB raw → 10–40MB as sorted zstd Parquet; NBBO quotes 5–10× that; [full depth](/guides/market-data-types/#the-ladder) 10–100× again. A 500-name universe's trades+quotes for a decade fits in low single-digit TB well-stored — one NVMe drive, [DuckDB-queryable](/guides/tick-data-databases/), no cluster. The cluster conversation starts with full-depth multi-venue data or [team concurrency](/guides/tick-data-databases/#defaults-wed-actually-pick), not before. ## Common ways tick storage fails - **The small-files explosion**: per-symbol-per-day files across thousands of symbols → filesystem metadata dominates, scans crawl. - **Float prices**: exact price-level queries (`px == 100.00`) silently miss; [order-book reconstruction](/guides/market-data-types/) breaks on representation error. - **Arrival-order writes**: unsorted files that compress badly and scan worse; sort at ingest, always. - **Timezone soup**: mixed local/UTC timestamps discovered during a [cross-venue study](/guides/evaluate-microstructure-paper/#the-empirical-checklist) — the archaeology nobody budgets. - **No ingest validation**: schema drift, duplicate days, and gap days enter silently; [checks at write time](/guides/notebook-to-production/) cost minutes and save weeks. ## Questions to ask (yourself, quarterly) 1. Can I scan a year of one symbol in seconds, and a day of all symbols in seconds? (If not, layout — not hardware.) 2. Is anything irreplaceable stored only in a derived format? 3. Would a [restore from backup](/guides/notebook-to-production/#6-backups-you-have-restored-at-least-once) reproduce byte-identical files? ## Where this connects This layout is what makes the [stack guide's](/guides/quant-research-stack/) Parquet+DuckDB default actually fast, what the [database engines](/guides/tick-data-databases/) ingest when you outgrow files, and the substrate the [PIT architecture](/guides/point-in-time-data-architecture/) builds its event-sourced layers on. --- **Engines on top:** [TimescaleDB vs ClickHouse vs DuckDB →](/guides/tick-data-databases/) · [Parquet vs Postgres →](/guides/parquet-vs-database/) · **The correctness layer:** [PIT architecture →](/guides/point-in-time-data-architecture/) · **Size it:** [tick data storage sizer →](/tools/tick-data-storage-sizer/) ---------------------------------------------------------------------- # How to Tell Whether a Trading Backtest Is Real URL: https://thequant.space/guides/evaluate-trading-backtest/ Section: Guides Date: 2026-09-07 Description: A practical framework for evaluating trading backtests: the eight questions that expose fake performance, the metrics that must be present, and the verdict rules we apply to published research. Most impressive backtests are wrong, and most are wrong *unintentionally* — which is worse, because the author's confidence is genuine. After scoring thousands of empirical papers for [rigor](/score-guide/), the failures cluster into eight questions. Ask them in order; each has a verdict rule. A backtest that survives all eight isn't proven — but one that fails two or more is not evidence of anything. ## 1. Could the strategy have known its own universe? Survivorship is the first check because it silently manufactures alpha before any signal is computed. The universe at each historical date must be "securities tradeable *on that date*," including everything that later delisted, defaulted, or was absorbed. **Verdict rule:** if the paper doesn't name a survivorship-clean source or describe delisting handling, assume 1–4%/yr of phantom performance on long equity strategies — enough to erase most published edges. Full treatment: [Survivorship bias →](/guides/survivorship-bias/) ## 2. Could the signal have known the future? Look-ahead leaks arrive through restated fundamentals, final index membership, same-bar execution, and — increasingly — ML models trained on text corpora that include the test period. **Verdict rule:** delay every signal by one bar and rerun. Alpha that halves under a one-bar lag was living on timing leakage; alpha that vanishes was never there. Full taxonomy: [Look-ahead bias & point-in-time data →](/guides/look-ahead-bias-point-in-time-data/) ## 3. What happens at 2× the assumed costs? Every backtest embeds a cost assumption, often zero. The question isn't whether the assumption is right — it's how fragile the result is to being wrong. **Verdict rule:** demand (or rebuild) the cost-sensitivity curve at 1×, 2×, 3× assumed costs, as required by [Gate 1 of our production checklist](/guides/production-checklist/#gate-1--validation-before-writing-any-production-code). High-turnover strategies that die at 2× costs are cost models wearing a strategy costume. The cost arithmetic itself: [Transaction costs & market impact →](/guides/transaction-costs-slippage-market-impact/) ## 4. How many independent bets is this, really? Four hundred trades across twelve correlated event clusters is twelve observations. Fifty markets on one election is one bet ([the prediction-market version](/guides/prediction-market-backtesting/#4-correlated-event-clusters)). Overlapping multi-day holding periods shrink effective sample size the same way. **Verdict rule:** cluster trades by underlying event/period before computing any significance; if the paper reports per-trade statistics on overlapping positions without adjustment, discount its t-stats heavily. Then apply the [multiple-testing lens](/guides/statistical-vs-economic-significance/) — how many variants were tried to find this one? ## 5. Does it survive the neighborhood of its own parameters? A real edge is a plateau; an artifact is a spike. Lookback 20 works, lookbacks 15 and 25 collapse → you're looking at noise that happened to fit. **Verdict rule:** ask for (or grid) the parameter neighborhood. Papers that report a single parameterization with no sensitivity analysis earn a "Lab Rats"-adjacent discount on our [rigor axis](/score-guide/) regardless of how good the single number looks. ## 6. Was out-of-sample actually out of sample? The phrase "out-of-sample" does heavy, often dishonest work. A holdout that was evaluated repeatedly during development is training data with extra steps. **Verdict rule:** look for the *protocol*, not the split — walk-forward structure, embargo periods, and evidence the holdout was touched once. The strongest tell of honesty is a paper reporting *worse* OOS than in-sample numbers and discussing why. Full protocol: [Walk-forward & out-of-sample testing →](/guides/walk-forward-out-of-sample-testing/) ## 7. Would the market let you do this at size? A 3% edge on a universe trading $200k/day is a hobby, not a strategy. **Verdict rule:** multiply implied position sizes against real ADV and book depth; anything above ~1–5% of daily volume needs an impact model, not a fee assumption. Papers on small caps, crypto long-tails, and prediction markets fail here constantly — the [phantom-liquidity problem](/guides/prediction-market-backtesting/#2-phantom-liquidity). ## 8. Is the reporting complete enough to be falsifiable? The minimum honest panel: returns **and** turnover **and** max drawdown **and** exposure over time **and** subperiod breakdown **and** the number of strategies tried. Each omission hides a specific failure mode (no turnover → cost fragility; no subperiods → one-regime wonder; no trial count → [p-hacking](/guides/statistical-vs-economic-significance/)). **Verdict rule:** treat missing panels as adverse inference — authors publish what flatters. ## The meta-test: luck has a distribution Any single equity curve — however smooth — is one draw from a fan of possible paths. Before believing a backtest, ask what the *fan* looks like: our [equity-curve simulator](/tools/equity-curve-simulator/) shows how wide the outcome distribution is for a strategy with a genuinely fixed edge, and the [deflated Sharpe ratio](/guides/deflated-sharpe-ratio/) quantifies how impressive the best of N attempts should look *by luck alone*. A backtest is real when it clears that bar — not when it merely looks good. ## The 60-second version | Question | Kill criterion | |---|---| | Universe | No survivorship handling stated | | Timing | Alpha dies at one-bar lag | | Costs | Dead at 2× assumed costs | | Sample | Correlated bets counted as independent | | Parameters | Spike, not plateau | | OOS | No protocol, or reused holdout | | Capacity | Size > ~5% ADV with no impact model | | Reporting | Missing turnover/drawdown/trials | --- **Apply it:** [highest-rigor papers](/) pass most of this by construction · [reproduce one yourself →](/guides/reproduce-quant-research/) ---------------------------------------------------------------------- # How to Version Datasets and Backtests URL: https://thequant.space/guides/versioning-datasets-backtests/ Section: Guides Date: 2026-09-07 Description: Versioning for quant research: content-addressed datasets, config-hashed backtests, the lineage chain that makes any result re-derivable, and the lightweight tooling that suffices. Six months from now, a [result row](/guides/quant-research-pipeline/#4-backtest--a-harness-not-a-script) says Sharpe 1.4 and you need to know: *exactly what produced this number?* If the answer requires memory, the result is folklore. Versioning is the discipline that makes every number re-derivable from **code version × config version × data version** — and in research (unlike app development) the data axis is the one that bites, because [data changes under you](/guides/corporate-actions-adjusted-prices/#the-mechanics-of-adjustment): vendors restate, adjustments rewrite history, and your own [ingestion pipeline evolves](/guides/point-in-time-data-architecture/#the-migration-path-from-a-naive-store). ## Versioning data: snapshots + content addresses The workable pattern for [file-tier data](/guides/parquet-vs-database/#the-two-tier-default), without heavy tooling: - **Immutability by convention**: partitions, once written, are never edited — corrections write *new* partitions ([the raw zone's append-only rule](/guides/quant-research-stack/#layer-1-market-data--own-your-pipeline-rent-the-feed) extended downstream). Mutation is the enemy of versioning; remove it and versioning reduces to naming. - **Content-address the snapshots**: a manifest per dataset version — file list + sizes + checksums (a few lines of hashing) — stored in git next to the code. `data_version = manifest hash`. Now "the data changed" is a diff, not a suspicion. - **Derived data versions by recipe**: features and bars are functions of (upstream version, transform code version); store *that pair* as their version and treat the outputs as [regenerable cache](/guides/store-tick-data-efficiently/#the-reference-layout). If regeneration doesn't reproduce byte-identical outputs, you've found nondeterminism worth killing. - **The database tier** versions differently: [bitemporal tables version themselves](/guides/point-in-time-data-architecture/#the-core-pattern-two-timestamps-never-one) (knowledge time *is* a version axis); for the rest, timestamped `pg_dump`s [you've actually restored](/guides/notebook-to-production/#6-backups-you-have-restored-at-least-once). Purpose-built tools (DVC, lakeFS, Delta) formalize all this; adopt one when manifests-in-git strains, not before — the concept, not the tool, is the requirement. ## Versioning backtests: the run record Every backtest run writes one row, at run time, automatically ([the harness's job](/guides/quant-research-pipeline/#4-backtest--a-harness-not-a-script), never a human's): ```text run_id | git_commit (code) | config_hash (full resolved params) | data_versions (manifest hashes consumed) | seed | environment (lockfile hash) | started/finished | metrics | artifacts_path ``` Three details carry the weight: **hash the *resolved* config** (defaults filled in — "the default changed" is the classic silent result-drift); **record data versions *consumed*, not "latest"** (the run must name its snapshots, or reruns quietly use different data); and **store the row even for failed and abandoned runs** — the run table is also your [trial counter](/guides/deflated-sharpe-ratio/#using-it-honestly-without-ceremony), and dead runs count. ## The lineage chain, end to end The standard a [reproduction](/guides/reproduce-quant-research/) of your *own* work should meet: any reported number traces to a run row → which names code commit, config hash, data manifests, seed, environment → each of which is fetchable → and rerunning reproduces the number to [stated tolerance](/guides/reproduce-quant-research/#stage-4--reconcile-match-statistics-before-results). That chain is what "reproducible" *means* operationally — and it's perhaps two days of harness work, mostly bookkeeping. The payoff compounds: [live-vs-backtest divergence](/guides/production-checklist/#gate-5--monitoring-and-reconciliation) debugging starts from "which exact inputs produced the research claim?", answered in seconds. ## Common versioning failures - **"Latest" as a data reference**: the run that can never be rerun; every downstream comparison contaminated by data drift. - **Notebook results with no run row**: [the exploratory finding that can't be re-derived](/guides/quant-research-stack/#layer-3-research-environment--optimize-for-iteration-speed) — the moment a number matters, it goes through the harness. - **Config in code**: parameters edited in place between runs; the git history *is* the config history only if configs live in files. - **Un-versioned environment**: dependency drift changes results months later; the [lockfile hash](/guides/docker-reproducible-research/) belongs in the run row. - **Adjustment-vintage amnesia**: results computed on last year's [adjusted prices](/guides/corporate-actions-adjusted-prices/) compared against this year's — the data-version axis's sneakiest instance. ## Questions to ask (of your own shop, quarterly) 1. Pick a number from three months ago: can you re-derive it, today, from its row? 2. Do any pipelines read "latest"? 3. Does the trial count in [DSR calculations](/guides/deflated-sharpe-ratio/) come from the run table or from memory? 4. Would a new teammate's rerun match yours byte-for-byte? ## What our scoring captures Papers can't show their run tables, but the textual shadows of this discipline — stated seeds, pinned versions, config listings, [code that reproduces specific tables](/guides/find-code-finance-papers/#the-15-minute-audit-before-found-code-touches-your-research) — are strong [rigor](/score-guide/) signals, and their absence in ML-heavy work is the genre's [reproducibility gap](/guides/evaluate-ml-trading-paper/) made visible. Internally, this guide is the difference between research and anecdote generation. --- **The process around it:** [research pipeline →](/guides/quant-research-pipeline/) · **The environment axis:** [Docker reproducibility →](/guides/docker-reproducible-research/) · **The experiment layer:** [ML tracking stack →](/guides/ml-experiment-tracking/) ---------------------------------------------------------------------- # Limit-Order-Book Imbalance: What It Measures and What It Misses URL: https://thequant.space/guides/limit-order-book-imbalance/ Section: Guides Date: 2026-09-07 Description: Order-book imbalance as a predictor: why it works at tick horizons, the spoofing and iceberg problems, the monetization gap, and how to evaluate imbalance-based research. **Not an implementation claim.** Imbalance's predictive power is real, replicated, and largely priced by the participants fast enough to use it; this guide teaches what the measure *is* so you can read [the LOB literature](/guides/map-lob-prediction/) without inheriting its optimism. **Definition.** Order-book imbalance at its simplest: I = (B − A)/(B + A), resting bid size minus ask size over their sum, at the top level or aggregated over depth. I > 0 says buying interest outweighs selling interest *among displayed passive orders* — and the next mid-move is more likely up than down. That conditional probability, often 60–75% at the next-tick horizon in liquid books, is the most replicated fact in [microstructure](/topics/market-microstructure/). ## Why it works: the mechanical channel The prediction is less mysterious than it looks — partly *mechanics*, not information: if the ask queue is thin, a modest market buy exhausts it and the mid moves up by construction; imbalance partially measures *how far the next trade will push*. Add a genuine informational channel (quoters skew away from anticipated moves, so their nets encode short-horizon views) and you get a strong, fast-decaying signal: predictive power concentrates at horizons of ticks-to-seconds and [decays toward noise](/guides/map-lob-prediction/) at minutes. The deep-learning LOB literature is, in large part, the project of squeezing more from this same object — with the recurring finding that fancy architectures beat linear imbalance features by margins that shrink [under honest evaluation](/guides/evaluate-ml-trading-paper/). ## What it misses — the four blind spots 1. **Hidden liquidity**: icebergs and fully hidden orders mean displayed size ≠ available size; measured imbalance can invert the true state exactly when someone large is working. 2. **Strategic display**: displayed size is a *choice*, and one manipulable at low cost — spoofing (illegal, extant) is precisely the practice of manufacturing imbalance signals; even legal quoting games make the displayed book a communication channel, not a census. 3. **Queue composition**: 10,000 shares of resting bid means different things if it's one institutional order vs fifty fast traders who'll cancel on the first adverse print. Imbalance counts shares; [toxicity lives in composition](/guides/market-making-mechanics/#adverse-selection-the-flow-you-dont-want-is-the-flow-you-get). 4. **Venue fragmentation**: one venue's book is a sample of the market; consolidated imbalance requires feeds most studies (and traders) don't have, and cross-venue imbalance dynamics differ from single-book ones. ## The monetization gap: the honest arithmetic The signal's paradox: a 70% next-tick edge sounds enormous and nets to nothing for most participants. The move being predicted is ~one tick or less; capturing it as a taker costs [spread + fees](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges) — usually more than a tick. As a maker, using imbalance means quoting *with* the signal — which is what every competing quoter does, so the edge expresses as [queue position and cancellation speed](/guides/market-making-mechanics/), i.e., infrastructure. This is why imbalance research is best read as **execution science, not alpha**: its uncontested production value is in timing child orders and improving fills — shaving cost per share — rather than standalone strategies. Papers translating imbalance accuracy into fee-free tick-capture backtests are the genre's standing fiction; the [horizon-overlap and label pathologies](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion) compound it. ## Common ways imbalance strategies fail in production - **The fee wall**: accuracy real, per-trade edge < round-trip cost — discovered only when the [cost columns](/guides/evaluate-trading-backtest/#3-what-happens-at-2-the-assumed-costs) are honest. - **Latency relativity**: the signal decays in milliseconds-to-seconds; being mid-pack fast means acting on stale imbalance against faster participants — negative selection, [ecology again](/guides/market-making-mechanics/#where-the-spread-goes-the-ecology-argument). - **Regime sensitivity of the mechanical channel**: tick-size and queue-length regimes change what imbalance means; a model trained on large-tick names [misfires on small-tick ones](/guides/evaluate-microstructure-paper/#the-worked-example-true-here-false-there). - **Feedback under deployment**: your own orders alter the book you're reading; at any size, the signal begins measuring you. ## Questions to ask an imbalance paper 1. What horizon, and what round-trip cost at that horizon? (The two numbers that decide everything.) 2. Are hidden-liquidity venues/instruments in the sample, and is robustness to icebergs discussed? 3. Is the evaluation [next-tick accuracy or costed P&L](/guides/evaluate-ml-trading-paper/#the-worked-example-the-accuracy-illusion) — and if P&L, how are fills simulated? 4. Single venue or consolidated — and does the claim respect which? ## What our scoring captures The [LOB-prediction map's](/guides/map-lob-prediction/) evaluation branch exists because this genre's rigor variance is extreme: identical architectures score anywhere on our [axis](/score-guide/) depending on label construction, cost handling, and horizon honesty. Imbalance papers reporting costed, latency-aware results are rare and rank accordingly; accuracy-only papers fill the hub's lower half. The measure is real; the score mostly tracks whether the *evaluation* was. --- **The context:** [market-making mechanics →](/guides/market-making-mechanics/) · [microstructure noise →](/guides/microstructure-noise-realized-volatility/) · **The field:** [LOB research map →](/guides/map-lob-prediction/) ---------------------------------------------------------------------- # Local GPU vs Cloud GPU for Financial NLP and LLM Research (2026) URL: https://thequant.space/guides/local-vs-cloud-gpu/ Section: Guides Date: 2026-09-06 Description: When to buy a GPU and when to rent one for financial NLP and LLM research: break-even math, VRAM sizing, data-licensing constraints, and the hybrid default. The GPU question generates more opinion per dollar than any other infrastructure decision in quant research. Strip the tribalism and it's arithmetic plus two constraints specific to finance — one legal, one about workload shape. This guide gives you the formula, the finance-specific gotchas, and the setup that wins for most research shops. Run your own numbers in our [compute-cost calculator](/tools/compute-cost-calculator/). ## Start with the workload, not the hardware Financial NLP and LLM research decomposes into four workload shapes with very different economics: 1. **Embedding and classification at scale** — encoding years of filings, news, or transcripts with sentence transformers or small finetuned models. Bursty: enormous for a week, then nothing. 2. **Fine-tuning small/mid models** (1–8B) on financial text. Bursty and checkpointable. 3. **Batch inference with LLMs** — scoring documents, extracting structured signals daily. Steady, small, latency-irrelevant. 4. **Serving/interactive work** — a research assistant, a dashboard. Steady, tiny. The pattern: research GPU demand is **spiky**. You need 500 GPU-hours one month and 20 the next. That asymmetry is the whole case for renting — and the case for owning is what happens between the spikes: the experiments you *don't* run when every hour has a visible price tag. People who own GPUs iterate more. That behavioral effect is real and belongs in the decision. ## The break-even arithmetic The formula that settles most debates: ``` monthly_ownership_cost = hardware_price / amortization_months + watts × hours_used / 1000 × $/kWh break_even_hours = monthly_ownership_cost / cloud_$_per_hour ``` Worked example with September 2026 magnitudes (all editable in the [calculator](/tools/compute-cost-calculator/) — verify current prices): - A capable local card in the 24GB class (RTX 4090-tier, used or new-ish) lands somewhere around **$1,500–2,500** with the PSU/case upgrades to feed it. Amortize over 30 months → **$50–85/month** plus electricity (~450W under load; at $0.10–0.30/kWh and 100 hours/month, another **$5–14**). - Community/marketplace GPU clouds (Vast.ai, RunPod, Lambda and peers) rent the same 24GB class for roughly **$0.25–0.60/hour**, and 80GB A100/H100 class for **$1–3.50/hour** depending on tier and interruptibility. Hyperscaler on-demand pricing runs 2–5× the marketplace rate. At $0.40/hour rental, a $2,000 card amortized over 30 months breaks even around **170–210 GPU-hours/month**. Below that, renting wins on cash; above it, owning wins — before the two finance-specific factors below, which move the line in both directions. ## The two constraints generic guides miss **1. Data licensing may decide for you.** Most commercial market-data and news agreements (and effectively all bulk news/transcript licenses) restrict redistribution and often restrict processing on infrastructure you don't control. Uploading licensed tick data or a news archive to a marketplace GPU host whose terms let them access attached storage can put you in breach — and marketplace clouds are exactly the cheap tier. Check your license *before* the cost spreadsheet; if it forbids third-party processing, your options are local hardware, a reserved single-tenant instance at several times the marketplace price, or the vendor's own cloud offering. This single clause makes local GPUs rational for shops whose raw economics say rent. **2. Financial text workloads are checkpoint-friendly.** Embedding runs and fine-tunes resume cleanly from checkpoints, which means you can use **interruptible/spot instances** — the cheapest tier of cloud — with little pain. Strategies that make cloud look expensive usually assume on-demand pricing; disciplined checkpointing cuts the effective rate by 2–3×. If your pipeline can't survive an interruption, fix that before buying hardware to work around it. ## VRAM sizing for LLM work (the only spec that matters first) Rules of thumb that hold in 2026: - **Embedding/classifier models** (up to ~1B): any 12GB card. This tier is cheap to own; never rent for it. - **7–8B models**: comfortable on 24GB at fp16; fine on 16GB quantized (Q5+ with headroom for context). LoRA fine-tuning of 7–8B fits on 24GB. - **13–14B**: 24GB quantized for inference; fine-tuning wants 48GB+. - **70B-class**: inference needs ~40GB quantized aggressively, 2×24GB with tensor splitting, or an 80GB card; fine-tuning is multi-GPU cloud territory, full stop. - **Frontier API models**: for many extraction/scoring tasks, a per-token API call beats both options — batch APIs price at roughly half of interactive rates, and for daily-scale document scoring the monthly bill is often single-digit dollars. Don't buy silicon to do what $8/month of API calls does better. (Our own [paper-scoring pipeline](/about/) runs this way.) The honest sizing question is: what's the largest model that measurably improves *your* signal? In financial NLP, the gap between a well-finetuned 8B and a 70B on classification/extraction tasks is frequently smaller than the gap between either and better labels. ## The hybrid default For a one-to-three-person research shop, the configuration that keeps winning: 1. **Own one 24GB-class card** in your workstation. It covers all embedding work, quantized inference up to 14B, LoRA fine-tunes of 8B, and — critically — unlimited zero-guilt experimentation. 2. **Rent interruptible big-GPU capacity** for the occasional 70B experiment or large sweep, with checkpointing as a hard requirement. Develop locally on subsampled data first; rent only to scale. 3. **Use batch APIs** for steady document-scoring pipelines where licensing permits. 4. Revisit ownership when your rental bill exceeds ~$80–100/month for three consecutive months — that's the amortized cost of the next card up. ## Decision table | Your situation | Answer | |---|---| | License forbids third-party processing | Local (or vendor cloud), regardless of cost math | | Bursty big-model experiments, checkpointable | Rent interruptible | | Steady daily batch scoring | Batch API first, local second | | < ~150 GPU-hours/month, no license constraint | Rent | | > ~200 GPU-hours/month sustained | Own (and still rent spikes) | | "I'll experiment more if it feels free" is true of you | Own the mid-tier card; it pays for itself in iterations | *Prices are September 2026 orders of magnitude — verify before purchase; GPU markets reprice monthly. No sponsored placements; if that changes it will be disclosed inline.* --- **Run your numbers:** [Compute-cost calculator →](/tools/compute-cost-calculator/) · **The rest of the stack:** [Quant research stack guide →](/guides/quant-research-stack/) ---------------------------------------------------------------------- # Look-Ahead Bias and Point-in-Time Data: The Complete Taxonomy URL: https://thequant.space/guides/look-ahead-bias-point-in-time-data/ Section: Guides Date: 2026-09-07 Description: Every way future information leaks into backtests: restated fundamentals, index membership, same-bar execution, timestamp semantics, and LLM training-data leakage — with detection tests for each. Look-ahead bias is knowing the future without noticing. Nobody codes `if date > today: peek()` — the future arrives through data that was *revised after the fact*, timestamps that mean something other than assumed, and models trained on text written later. It is the second member of the bias family after [survivorship](/guides/survivorship-bias/), harder to spot because each leak has its own disguise. Here is the full taxonomy, each entry with its detection test. ## 1. Restated fundamentals — the classic Companies restate earnings; standard data feeds silently overwrite history with the corrected numbers. Backtest on restated data and your value signal in March is using figures the market didn't see until August. **Point-in-time (PIT)** data — as-reported values with their announcement dates and the full restatement trail — is the fix, and its availability is the sharpest quality line between [data-vendor tiers](/guides/market-data-vendors/#2-is-fundamental-data-point-in-time). *Detection:* pick a company with a known restatement; check whether your dataset shows the original or corrected figure for the original period. Also watch the tell in results: fundamental signals with implausibly *fast* reaction to information. ## 2. Reporting-lag compression Even unrestated fundamentals leak if used at period-end rather than announcement date: Q4 earnings dated December 31 weren't public until February. A signal formed "monthly" from quarter-end data is trading weeks before knowledge existed. *Detection:* lag all fundamental inputs to announcement date (or a conservative fixed lag) and rerun — the standard robustness check the best [factor papers](/topics/factor-investing/) report by default. ## 3. Retroactive universe and membership "S&P 500 since 2010, current membership" imports two biases at once — [survivorship](/guides/survivorship-bias/), plus the look-ahead of knowing *who would be added*: index inclusion itself moves prices, so a today's-members universe pre-selects stocks before their inclusion pop. *Detection:* rebuild with as-of-date constituents; strategies that shrink meaningfully were harvesting membership foresight. ## 4. Same-bar execution Signal computed on today's close, filled at today's close: information and execution occupy the same instant. Related: signals built on daily *high/low* (you can't know the day's high until it's over), and stops filled at exact stop prices through gaps. *Detection:* the **one-bar-lag test** from the [backtest framework](/guides/evaluate-trading-backtest/#2-could-the-signal-have-known-the-future) — execute at the next bar's open; collapse means the edge was timing leakage. Intraday mean-reversion backtests are the serial offenders. ## 5. Timestamp semantics The subtle sibling of #4: fields that don't mean what you assume. Is the daily "close" the 4:00pm auction or the last extended-hours tick? Is the bar stamped at open or close of its interval? Exchange time or vendor receipt time? Which timezone, and is it DST-stable? Our [Monte Carlo audit](/guides/auditing-a-monte-carlo/#11-verify-field-semantics-against-a-finer-dataset) shows the technique: **cross-validate the coarse dataset against a finer one you can reconcile** — one minute-bar request settles what documentation leaves open. ## 6. Normalization and statistics computed on the full sample Z-scoring features with the whole sample's mean and variance, selecting features by full-sample correlation, or scaling by full-period volatility — each quietly hands early observations knowledge of the future distribution. The ML-pipeline version of look-ahead, endemic in the [machine-learning literature](/topics/machine-learning/). *Detection:* every transform must be fit on data available at prediction time — expanding or rolling windows only. If the pipeline calls `fit` anywhere on the full dataset before splitting, it's leaking. Related boundary leakage — overlapping windows straddling the train/test split — is covered in [walk-forward testing](/guides/walk-forward-out-of-sample-testing/). ## 7. LLM training-data leakage — the newest member A 2024-trained language model "predicting" 2023 headlines has read the endings. Any pipeline using pretrained embeddings or LLM scores must ensure the model's **training cutoff predates the test period** — or restrict claims to tasks (extraction, classification quality) where knowing the future doesn't help. Many published LLM-trading backtests fail exactly this test; it's the first thing we check in the [NLP & LLMs hub](/topics/nlp-llm/). *Detection:* compare performance before vs after the model's knowledge cutoff — a cliff at the cutoff date is the confession. ## The point-in-time discipline, positively stated The cure across the taxonomy is one rule: **every input must carry the timestamp at which it became knowable, and the backtest may only read inputs whose knowable-time precedes decision-time.** In practice: PIT fundamentals with announcement dates; as-of-date universes; next-bar execution; expanding-window transforms; cutoff-checked models; and an event-log data architecture where revisions append rather than overwrite — which is exactly how the [raw zone](/guides/quant-research-stack/#layer-1-market-data--own-your-pipeline-rent-the-feed) and an append-only [research database](/guides/tick-data-databases/) are designed. The economics are asymmetric: PIT-clean data costs more and yields *smaller* backtested returns — which is precisely why it's worth it. You're not paying for better numbers; you're paying to stop lying to yourself. --- **The bias family:** [Survivorship →](/guides/survivorship-bias/) · [Walk-forward & OOS →](/guides/walk-forward-out-of-sample-testing/) · **Data that supports PIT:** [vendor guide →](/guides/market-data-vendors/) ---------------------------------------------------------------------- # Market Data for Quant Research: How to Choose a Vendor (2026) URL: https://thequant.space/guides/market-data-vendors/ Section: Guides Date: 2026-09-06 Description: A decision framework for choosing market data vendors for quant research: survivorship bias, point-in-time integrity, licensing, and the real cost tiers. Market data is the one purchase every systematic trader makes, and the one where marketing pages are least informative. Vendors compete on symbol counts and API polish; backtests die from properties that never appear on a pricing page — survivorship handling, restatement history, timestamp semantics. This guide gives you the checklist that actually predicts research quality, then maps the 2026 landscape by tier. **Disclosure: this guide currently contains no affiliate links or sponsored placements. If that changes, links will be labeled inline.** ## The five questions that matter more than price ### 1. Does history include the dead? **Survivorship bias** is the classic retail-backtest killer: a universe of *currently listed* symbols excludes every bankruptcy and delisting, inflating long-strategy returns by roughly 1–4% annually on US equities depending on period and universe (the effect is largest in small caps). Test any vendor before trusting them: pick five well-known delistings and query their full history. If the data isn't there — or worse, is silently missing final months of trading — the vendor is for dashboards, not research. ### 2. Is fundamental data point-in-time? Companies restate earnings. A **point-in-time (PIT)** feed gives you the numbers *as investors saw them on each date*, with the restatement trail; a non-PIT feed silently overwrites history with corrected figures. Backtesting on restated data is look-ahead bias with a lag of months — value and quality factors are especially distorted. PIT fundamentals are the sharpest quality line in the market: they're what separates the institutional tier (and a handful of serious mid-tier vendors) from the commodity APIs. If a cheap vendor claims PIT, ask for the as-reported/restated field pair and spot-check a known restatement. ### 3. What do the timestamps mean? For anything intraday, ask: exchange timestamp or vendor receipt time? Which timezone, and is it DST-consistent? Are daily bars stamped at the exchange close or midnight UTC? Consolidated tape or single venue? None of this is exotic — but every one of these has silently shifted a backtest's signals by a bar. The tell of a good vendor is documentation that answers these without a support ticket. ### 4. Corporate actions: who does the adjusting? Splits and dividends must be handled *somewhere*. The clean setup is **unadjusted prices plus a separate corporate-actions table**, letting you adjust reproducibly in code. Pre-adjusted-only feeds are acceptable for quick research but make cross-vendor validation and dividend-strategy work painful. Check dividend treatment specifically: price-only adjustment understates total returns by the dividend yield, compounded. ### 5. What does the license actually permit? Personal-use plans commonly prohibit redistribution, publishing derived datasets, or commercial use — which can include managing outside money. Exchange fees for *real-time* data are a separate layer with their own professional/non-professional distinction; historical and delayed data is far less encumbered. Read the license before building a business on a feed, and budget time for this if you ever take outside capital. ## The 2026 vendor landscape, by tier Prices are order-of-magnitude as of September 2026 — verify current terms; this market reprices constantly. ### Commodity API tier (~$0–100/month) Alpha Vantage, Finnhub, Twelve Data, EODHD, Tiingo and similar. Fine for prototyping, dashboards, and learning. The common failure modes: thin or absent delisted history, non-PIT fundamentals, and rate limits that make universe-wide downloads slow. Tiingo and EODHD punch above their price on data hygiene; verify survivorship handling for your specific universe regardless. ### Serious individual tier (~$50–500/month) - **Polygon.io** — the default US equities/options API for individual quants: full-market coverage, real tick history, websockets, sane licensing. Check delisted coverage depth against your backtest start date. - **Databento** — usage-based pricing, exchange-native feeds (including futures via CME), precise timestamping, and the best documentation-honesty in the tier; the pay-per-query model rewards the raw-zone archiving habit from our [stack guide](/guides/quant-research-stack/). - **Norgate** — daily-frequency specialist for US/AU equities with meticulous delisting and corporate-action handling; a quiet favorite for end-of-day systematic research. - **Crypto:** exchange APIs are free and canonical for their own venues; aggregators (Kaiko, CoinAPI, Amberdata) matter only when you need normalized multi-venue history. ### Institutional tier (four figures/month and up) CRSP (the academic gold standard for survivorship-clean US equities), Compustat/S&P for PIT fundamentals, LSEG/Refinitiv, Bloomberg, FactSet, and direct exchange feeds. You buy this tier when compliance, PIT rigor, or asset-class breadth demands it — not for API convenience. Academic affiliations get CRSP/Compustat access through WRDS at a fraction of commercial cost; it remains the cheapest legitimate path to research-grade equity history. ## A decision framework in four steps 1. **Write down your universe and horizon first.** "US equities, daily, 15 years, long/short" and "BTC perpetuals, 1-minute, 3 years" have almost disjoint vendor shortlists. Most vendor regret comes from buying breadth you never query. 2. **Run the survivorship test** (five dead tickers) and the **timestamp test** (one known market event, checked bar by bar) on a trial plan before subscribing annually. 3. **Buy prices cheap, fundamentals carefully.** The tiers above apply to prices; for fundamentals, the PIT question dominates everything else including price. 4. **Archive everything you pull** (raw, append-only). Vendor churn is normal over a research career; your archive is what makes it survivable. ## Red flags, condensed - "10,000+ tickers!" with no mention of delistings → dashboard data. - Fundamentals with no as-reported/restated distinction → look-ahead bias on tap. - Pricing page without a license summary → licensing surprise later. - No documented timestamp semantics for intraday data → debugging session prepaid. --- **Next:** [Which database should hold all this? TimescaleDB vs ClickHouse vs DuckDB vs kdb+ →](/guides/tick-data-databases/) · **Budget it:** [tick data storage sizer →](/tools/tick-data-storage-sizer/) ---------------------------------------------------------------------- # Market Making: Inventory Risk, Adverse Selection, and Spread Capture URL: https://thequant.space/guides/market-making-mechanics/ Section: Guides Date: 2026-09-07 Description: The three forces of market making — spread capture, inventory risk, and adverse selection — with the Avellaneda-Stoikov intuition, the profitability identity, and the production failure catalog. **Not an implementation claim.** Market making is an infrastructure business wearing a strategy costume; this guide teaches the *forces* so you can evaluate [the research](/guides/map-market-making/) and recognize why most "market-making backtests" are fiction — not so you can quote tomorrow. **Definition.** A market maker posts simultaneous buy and sell quotes, earning the spread from uninformed flow while managing two costs: **inventory risk** (holding what you were hit on while the price moves) and **adverse selection** (being hit *because* the price is about to move). All of market-making research is the study of this triangle. ## The P&L identity that organizes everything Every market-making result decomposes as: ```text P&L = spread captured − adverse-selection losses − inventory-holding costs − fees + rebates ``` Measure the terms separately or understand nothing: spread capture scales with *uninformed* volume share; adverse selection is measured by **markouts** (mid-price change N seconds/ticks after your fill — the field's core diagnostic); inventory cost scales with variance held × holding time. A strategy can improve gross capture while dying on markouts, and vice versa — single-number P&L hides which force moved, which is why [empirical MM papers](/guides/evaluate-microstructure-paper/) that report markout curves outrank those reporting only returns. ## Inventory: the Avellaneda–Stoikov intuition The stochastic-control lineage's core insight, worth having even without the math: **your reservation price is not the mid** — it shifts against your inventory (long inventory → quote lower on both sides, shading to sell), and the shift grows with volatility, risk aversion, and remaining horizon. Every practical quoting system embeds some version of this skewing, learned or derived; the [RL market-making branch](/guides/map-market-making/) mostly rediscovers it with extra parameters. The model's known counterfactuals — Poisson arrivals, no informed flow — are exactly [the assumption-audit targets](/guides/evaluate-microstructure-paper/#the-theory-checklist) when reading descendants. ## Adverse selection: the flow you don't want is the flow you get Quotes are free options written to the market, and informed traders exercise them. Mechanics that follow: **queue position is value** (in price-time priority, front-of-queue fills contain less adverse selection — back-of-queue fills at the same price are the toxic residue); **cancellation speed is defense** (the spread compensates *average* toxicity, so slower reactions than the venue's ecology means systematic negative markouts); and **spread width is a filter setting** (wider quotes select for impatient/uninformed flow but capture less volume — the trade-off *is* the business). The measurement discipline: bucket markouts by counterparty urgency proxies (trade size, sweep vs single, time-of-day) — toxicity is not uniform, and neither should quotes be. ## Where the spread goes: the ecology argument Displayed spread ≠ your margin. It funds, in order: adverse selection to faster/better-informed participants, exchange fees net of rebates, inventory variance, and *then* profit. The ecology implication: market-making profitability is **relative** — to the venue's other quoters' speed and models — not absolute. This is why [MM research generalizes worse than anything else in microstructure](/guides/evaluate-microstructure-paper/#the-worked-example-true-here-false-there): a 2014 profitability result is a statement about 2014's ecology. AMM/DeFi research replays the same triangle with the quoting function fixed by the pool curve — impermanent loss *is* adverse selection in continuous form, LP fees *are* the spread — which is why [that branch](/guides/map-market-making/) reads as classical MM in new notation. ## Common ways market making fails in production - **The backtest quoted where the book wouldn't have let it**: simulated fills ignore queue position; real fills at that price would have been the back-of-queue toxic ones. Fill simulation without book reconstruction is the genre's [label-as-return equivalent](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion). - **Inventory death spiral in trends**: skewing reduces but doesn't eliminate accumulation against a trend; without hard position limits and a [kill switch](/guides/production-checklist/#gate-4--execution-safety-the-account-savers), one directional day returns a month of spreads. - **Toxicity regime shifts**: a new fast participant arrives; markouts degrade for weeks before cumulative P&L makes it obvious. Monitor markouts daily, [not P&L](/guides/regime-dependence/#the-live-trading-diagnosis-problem). - **Rebate-dependency**: strategies profitable only via maker rebates are fee-schedule bets; venue pricing changes are their regime breaks. - **The uptime tax**: quoting obligations and technology failures interact — being *down* during the volatile hour concentrates your uptime into exactly the adverse periods. ## Questions to ask an MM paper 1. Are markouts reported, at multiple horizons, by flow type? 2. How are fills simulated — and is queue position modeled at all? 3. Which venue/era ecology produced the data, and does the paper claim beyond it? 4. For RL variants: does the learned policy's skewing behavior [reduce to A-S intuitions](/guides/evaluate-rl-trading-paper/#questions-to-ask-while-reading), and is the environment's flow toxicity calibrated to anything real? ## What our scoring captures The [map of MM research](/guides/map-market-making/) sorts the field's branches; within them, the rigor axis rewards markout-level measurement and book-reconstruction honesty — the two practices separating empirical MM work from spread-times-volume arithmetic. Theory branch papers land in [Lab Rats](/score-guide/) by design; the calibrated exceptions top the hub. --- **The book's information content:** [LOB imbalance →](/guides/limit-order-book-imbalance/) · **The noise floor:** [microstructure noise & realized vol →](/guides/microstructure-noise-realized-volatility/) · **The field map:** [market-making research →](/guides/map-market-making/) ---------------------------------------------------------------------- # Microstructure Noise and Realized Volatility URL: https://thequant.space/guides/microstructure-noise-realized-volatility/ Section: Guides Date: 2026-09-07 Description: Why high-frequency volatility estimates explode: bid-ask bounce, discreteness, the signature plot, noise-robust estimators, and the practical sampling rules for realized volatility. **Not an implementation claim.** This is measurement mechanics — the layer under every [volatility-forecasting result](/guides/map-volatility-forecasting/) and high-frequency backtest, and a classic place where naive pipelines silently produce garbage inputs. **Definition.** Realized volatility (RV) estimates variance by summing squared high-frequency returns: RV = Σ r²ᵢ over a day's intervals. In theory, finer sampling → better estimates. In practice, observed prices = efficient price + **microstructure noise** — bid-ask bounce (trades alternate between bid and ask, manufacturing back-and-forth "returns"), price discreteness, and asynchronous updates — and squared returns sum the noise along with the signal. Sample every second and noise *dominates*: RV diverges as sampling accelerates, the opposite of what statistics promises. ## The worked example: the signature plot The diagnostic every practitioner should run once: compute RV for the same day at many sampling frequencies (1s, 5s, 30s, 1m, 5m, 15m, 30m) and plot RV against frequency. Clean theory predicts a flat line; reality shows RV **exploding at high frequencies** — a liquid stock might show annualized vol of 15% at 15-minute sampling and 45% at 1-second sampling, the excess being pure bounce. The plot's flattening point is your instrument's noise horizon, and it varies enormously: seconds for ES futures, minutes for liquid stocks, tens of minutes for small caps and [long-tail crypto](/topics/crypto-defi/). One plot per instrument class replaces a semester of theory — the same [cross-validate-against-finer-data instinct](/guides/auditing-a-monte-carlo/#11-verify-field-semantics-against-a-finer-dataset) applied to your own estimator. ## The estimator toolkit, in practice order - **Sparse sampling (5-minute RV)**: the industry default — sample slow enough that noise is negligible relative to signal. Throws away data; robust; the [HAR-RV baseline](/guides/map-volatility-forecasting/) that's so hard to beat is built on it. - **Subsampled/averaged RV**: average multiple 5-minute grids at different offsets — recovers some discarded efficiency for free; the sensible default upgrade. - **Realized kernels / two-scale estimators**: the econometrically serious tools — noise-robust by construction, usable at high frequency, with tuning parameters that [need their own validation](/guides/robustness-trading-strategy/). - **Mid-quote construction**: computing returns from mid-quotes rather than trade prices removes bounce at the source (not discreteness or quote noise) — often the cheapest large improvement, [data permitting](/guides/market-data-types/). - **Pre-averaging and cleaning**: outlier/bounce-back filters on ticks before anything else; the unglamorous step that prevents [one bad print from owning your day's estimate](/guides/auditing-a-monte-carlo/). ## Why it matters beyond estimation Downstream contamination is the real stakes: **volatility forecasting** benchmarks are only comparable when noise handling matches — a "superior" forecaster fed cleaner RV inputs is [an unfair race](/guides/evaluate-trading-backtest/); **[volatility targeting](/guides/volatility-targeting/)** and risk models inherit estimator bias directly into position sizes; **high-frequency Sharpe ratios** computed on noisy returns overstate both risk and opportunity (bounce looks like tradeable mean reversion — the classic fake signal, and the null every [tick-level reversion claim](/guides/map-lob-prediction/) must beat); and **variance-based [option strategies](/topics/options-derivatives/)** settle against realized measures whose construction is contractually specified — mismatch your estimator and hedge the wrong quantity. ## Common ways this fails in production - **The fake reversion strategy**: backtest on trade prices "discovers" tick-level mean reversion that is bid-ask bounce; live trading pays the spread to harvest a signal that *is* the spread. - **Estimator regime breaks**: tick-size changes, decimalization-era differences, and venue migrations shift noise properties mid-sample — [era-tagging applies](/guides/evaluate-microstructure-paper/#the-worked-example-true-here-false-there). - **Crypto's dirty ticks**: wash trading and venue-quality variance make exchange choice part of the estimator; RV from a manipulated venue is a manipulated input. - **Silent frequency upgrades**: a pipeline moved from 5-minute to 1-minute bars "for more data" quietly doubles measured vol and halves every vol-scaled position. ## Questions to ask a paper using realized measures 1. What sampling frequency and estimator — and is a signature plot (or noise-robustness check) shown? 2. Trade prices or mid-quotes? 3. Are competing forecasts fed the *same* realized input? 4. For tick-level trading claims: does the effect survive on mid-quote returns? ## What our scoring captures Estimator hygiene is quietly one of the rigor axis's sharpest discriminators in the [volatility hub](/topics/volatility/): papers specifying sampling scheme, estimator, and cleaning rules cluster at the top; papers computing "volatility" from raw ticks without comment populate the bottom. The [map's realized-vol branch](/guides/map-volatility-forecasting/) is effectively ranked by this hygiene. --- **Downstream:** [volatility targeting →](/guides/volatility-targeting/) · [vol forecasting map →](/guides/map-volatility-forecasting/) · **The noise's source:** [market-making mechanics →](/guides/market-making-mechanics/) ---------------------------------------------------------------------- # OHLCV, Trade, Quote, and Order-Book Data: What Each Can Answer URL: https://thequant.space/guides/market-data-types/ Section: Guides Date: 2026-09-07 Description: The market-data hierarchy from daily bars to full order books: what each granularity can and cannot answer, the storage and cost jumps between levels, and matching data type to research question. Market data is a ladder, and every research question has a *minimum rung*: try to answer it from below and you get [artifacts, not findings](/guides/microstructure-noise-realized-volatility/#common-ways-this-fails-in-production). This guide maps rung → capabilities → cost, so the [data budget](/guides/quant-research-stack/#layer-1-market-data--own-your-pipeline-rent-the-feed) matches the question. ## The ladder **1. Daily OHLCV** (~KB/instrument/year). Answers: cross-sectional strategies at daily+ horizons, [factor research](/topics/factor-investing/), trend/[vol-targeting](/guides/volatility-targeting/) at portfolio level. Cannot answer: anything intraday, true [execution costs](/guides/transaction-costs-slippage-market-impact/), and even "daily volatility" beyond crude estimators. Hidden subtleties this rung still carries: [which close](/guides/look-ahead-bias-point-in-time-data/#5-timestamp-semantics), [which adjustments](/guides/corporate-actions-adjusted-prices/), and consolidated-vs-primary volume. **2. Intraday bars** (1s–1m; ~MB/instrument/year). Adds: [realized volatility](/guides/microstructure-noise-realized-volatility/) at sane sampling, intraday seasonality, execution-window analysis, [event studies at honest timestamps](/guides/event-driven-research/#problem-2-the-timestamp-is-the-methodology). Still hides: everything inside the bar — bars are *summaries*, and bar-construction choices (time vs tick vs volume bars) are themselves [methodology](/guides/evaluate-ml-trading-paper/). **3. Trade prints (tick data)** (~10–100MB/instrument/year liquid). Adds: actual transaction analysis, volume profiles, trade-size distributions, signature plots. The classic ceiling: trades alone can't separate buyer- from seller-initiated flow reliably (inference rules like Lee-Ready guess, with known error rates) — and can't see [the liquidity that didn't trade](/guides/limit-order-book-imbalance/#what-it-misses--the-four-blind-spots). **4. Quotes / top-of-book (L1)** (~100MB–1GB/instrument/year). Adds: spreads as they stood (true [cost measurement](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges)), mid-quote returns (killing bounce artifacts), trade classification against prevailing quotes, [basic imbalance](/guides/limit-order-book-imbalance/). This rung is where most *execution research* becomes possible — and where [timestamp discipline](/guides/evaluate-microstructure-paper/#the-empirical-checklist) becomes the binding constraint. **5. Order-book snapshots / depth (L2)** (~1–10GB/instrument/year). Adds: depth-aware [impact and capacity](/guides/capacity-constraints/) estimates, book-shape features, [fill simulation with queue approximations](/guides/market-making-mechanics/#common-ways-market-making-fails-in-production). Snapshot frequency matters: 1-second snapshots miss the intra-second life where [HFT ecology](/topics/high-frequency-trading/) lives. **6. Full event feeds (L3/MBO)** (~10–100GB+/instrument/year). Adds: exact book reconstruction, queue-position modeling, order-lifecycle analysis (cancellations, [spoofing-adjacent patterns](/guides/limit-order-book-imbalance/)), honest [market-making backtests](/guides/market-making-mechanics/). The research rung of record for [microstructure claims](/guides/evaluate-microstructure-paper/#the-empirical-checklist) — and the storage rung where [database choices](/guides/tick-data-databases/) stop being optional. ## The two classic mismatches **Working a rung too low** — the artifact factory: "mean reversion" from OHLCV that's [bid-ask bounce](/guides/microstructure-noise-realized-volatility/#the-worked-example-the-signature-plot); execution costs assumed because quotes weren't available; capacity claims with no depth data behind them. If the claim mentions spreads, fills, or impact, it needs rung 4+; if it mentions queues, rung 6. **Working a rung too high** — the budget bonfire: buying L3 feeds to research monthly factor tilts burns [storage](/guides/store-tick-data-efficiently/), money, and engineering time on precision the question can't use. The [notebook-to-production stack](/guides/notebook-to-production/) principle applies: buy the rung the *bottleneck* demands. ## The self-collection option For rungs 4–6, one under-used path: **record your own** — exchange websockets (crypto natively; equities via vendor feeds) into your [raw zone](/guides/quant-research-pipeline/#2-data--one-clean-layer-biases-handled-once). Storage is cheap; history is unbuyable. The [prediction-market rule](/guides/prediction-market-backtesting/#2-phantom-liquidity) generalizes: start recording *before* you need the history, because rung 5–6 historical data is the hardest to purchase after the fact. ## Questions to ask (of a paper, or your own plan) 1. What's the *minimum rung* for this claim — and is the paper on it? 2. Bars: which construction, and does the result survive an alternative? 3. Trades: how is direction classified, and what's the stated error rate? 4. Books: snapshots or events, at what frequency, single-venue or consolidated? ## What our scoring captures Rung-appropriateness is scored implicitly: [rigor](/score-guide/) rewards data descriptions precise enough to check, and the [microstructure hub's](/topics/market-microstructure/) ranking correlates strongly with rung honesty — L3-based measurement work at the top, OHLCV-based "microstructure" at the bottom. State your rung; the readers who matter check. --- **Storage for the upper rungs:** [tick data efficiently →](/guides/store-tick-data-efficiently/) · [database comparison →](/guides/tick-data-databases/) · **The budget:** [free vs paid →](/guides/free-vs-paid-market-data/) ---------------------------------------------------------------------- # Options-Implied Volatility: A Practical Research Primer URL: https://thequant.space/guides/options-implied-volatility-primer/ Section: Guides Date: 2026-09-07 Description: Implied volatility for researchers: what IV actually is, surface anatomy, the variance risk premium, using IV as a signal, and the data pitfalls that corrupt options research. **Not an implementation claim.** Options research has the field's [least accessible data](/guides/find-datasets-quant-papers/#the-standard-dataset-zoo) and some of its most persistent premia claims; this primer is the conceptual toolkit for *reading* that literature without the classic misreadings. **Definition.** Implied volatility is the σ that makes the Black-Scholes formula reproduce an option's market price — **a price re-expressed in volatility units**, not a forecast. Everything useful and everything misleading about IV follows from that: it aggregates the market's risk-neutral expectation *plus* every premium (variance risk, tail risk, supply-demand pressure) into one convention-laden number. ## Surface anatomy in four facts 1. **The smile/skew**: equity index IV rises for low strikes — OTM puts price crash insurance at a premium. Skew is therefore a *quantity of fear priced*, and its steepness a sentiment/positioning object in its own right. 2. **Term structure**: usually upward-sloping in calm (vol mean-reversion from low levels + risk premia), inverting in stress — the VIX curve's shape is a regime indicator [old enough to be crowded](/guides/regime-dependence/#detecting-regime-carried-results-in-a-paper). 3. **Moneyness conventions matter**: surfaces plotted in delta vs strike vs log-moneyness behave differently under moves; papers comparing "fixed" points on differently-parameterized surfaces are [comparing different objects](/guides/evaluate-microstructure-paper/#the-empirical-checklist). 4. **IV is model-shaped but convention-robust**: using B-S to *quote* doesn't assume B-S is true — but Greeks and surface interpolation inherit model assumptions, and [pricing-model papers](/topics/options-derivatives/) live or die on that distinction. ## The variance risk premium: the bias that is the point IV systematically exceeds subsequently realized volatility in index options — the **variance risk premium (VRP)**: option sellers earn, on average, for bearing variance risk, [smoothly until abruptly](/guides/high-sharpe-ratio-not-investable/#3-volatility-isnt-risk--skew-is). Three readings coexist: as *forecast bias* (IV is an upward-biased vol predictor — subtract the premium before using it as one), as *harvestable premium* (short-vol strategies — the canonical [negative-skew trade](/guides/alpha-beta-alternative-risk-premia/#common-ways-the-taxonomy-fails-its-users-in-production) whose Sharpe flatters until the event), and as *information* (VRP's own variation predicts returns in a large literature — with all the [regime caveats](/guides/regime-dependence/) attached). Papers conflating the readings — treating the premium as forecast error to be "corrected," or its harvest as alpha — are the genre's standing confusion. ## IV as research input: what survives scrutiny The defensible uses: **IV as conditioning variable** ([vol-forecasting](/guides/map-volatility-forecasting/) models that include IV beat history-only ones — the surface *does* contain forward information, premium and all); **skew/term-structure changes as flow proxies** (positioning shifts show in surface shape before realized vol); and **cross-sectional IV signals** (name-level IV anomalies live in the [factor literature's](/topics/factor-investing/) options wing, with its [multiple-testing baggage](/guides/statistical-vs-economic-significance/)). The suspicious use: strategies "trading the gap" between IV and a realized forecast [without pricing the premium they're shorting](/guides/evaluate-trading-backtest/) — most are VRP harvests with extra steps and extra fees. ## Data pitfalls that corrupt options research - **Stale and wide quotes**: away-from-the-money options trade rarely; "prices" are quotes, [mid-of-wide-spread constructions](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges), and IV computed from them inherits the noise — surface-fitting papers must state their filters (volume, spread, arbitrage checks), and the filters are [construction choices](/guides/evaluate-factor-investing-paper/#the-worked-example-construction-is-the-hidden-factor). - **Synchronization**: option close and underlying close are timestamped [minutes apart](/guides/look-ahead-bias-point-in-time-data/#5-timestamp-semantics); computed IVs at "the close" embed spurious movements. Serious studies use synchronized snapshots. - **American/dividend adjustments**: single-name IVs need early-exercise and dividend handling; sloppy versions manufacture skew artifacts. - **Survivorship of the surface**: strikes and expiries are listed *in response to* demand and price levels — the surface's coverage is [endogenous](/guides/survivorship-bias/), and backtests on "all listed options" inherit that selection. - **Costs at options scale**: spreads as a fraction of premium are enormous; options backtests gross of costs are [more fictional than equity ones](/guides/evaluate-trading-backtest/#3-what-happens-at-2-the-assumed-costs) by an order of magnitude. ## Questions to ask an options paper 1. Quote filters, synchronization, and exercise-style handling — stated? 2. Is the claimed effect a repackaged VRP harvest — what's the skew/kurtosis of its returns? 3. Costed at realistic option spreads, or at mid? 4. Does the surface parameterization drive the result? ## What our scoring captures The [options hub](/topics/options-derivatives/) is our most [Lab-Rats-heavy](/score-guide/) territory — pricing machinery scores high on math, and the rigor tier belongs to the minority doing synchronized-data, filtered-quote, costed empirics. The VRP conflations this guide flags are visible in scored papers as a pattern: high-Sharpe short-vol results whose rigor score is dragged by missing tail accounting — the axis working as intended. --- **The measurement floor:** [realized vol & noise →](/guides/microstructure-noise-realized-volatility/) · **The forecasting race:** [vol research map →](/guides/map-volatility-forecasting/) · **The premium's cousin:** [volatility targeting →](/guides/volatility-targeting/) ---------------------------------------------------------------------- # P-Hacking in Financial Research: The Practices, the Tells, the Fixes URL: https://thequant.space/guides/p-hacking-financial-research/ Section: Guides Date: 2026-09-07 Description: How p-hacking works in finance: the seven questionable research practices, the tells visible in published papers, and the reader-side corrections that keep you from funding other people's noise. **Definition.** P-hacking is the family of practices that manufacture statistical significance through flexibility: adjusting samples, specifications, and metrics *after seeing results* until p < 0.05 appears. In finance it is rarely fraud and usually drift — each choice locally defensible, the sequence jointly fatal. It's the human-behavior layer beneath [backtest overfitting](/guides/backtest-overfitting/) (the selection mechanics) and the [multiple-testing problem](/guides/statistical-vs-economic-significance/) (the field-level arithmetic); this guide is about recognizing the *practices* — in your own work and in the papers you read. ## The seven practices, in ascending order of self-deception 1. **Outcome switching** — the study set out to test monthly returns; weekly "turned out to be more appropriate." 2. **Sample sculpting** — dropping "anomalous" periods (2008, 2020, that one bad year), filtering "illiquid" names *after* seeing they hurt, or the start date that [happens to flatter](/guides/how-to-read-quant-finance-papers/#red-flags-that-end-a-read-early). 3. **Specification shopping** — trying controls, weightings, and breakpoints until significance; reporting the winner as *the* specification ([the factor-paper version](/guides/evaluate-factor-investing-paper/#the-worked-example-construction-is-the-hidden-factor)). 4. **Covert stopping rules** — extending or truncating the sample until the result stabilizes where wanted. 5. **Metric shopping** — Sharpe didn't clear the bar, but Sortino did; accuracy failed, but F1 passed ([the ML variant](/guides/evaluate-ml-trading-paper/)). 6. **Subgroup mining** — the effect "concentrates" in small caps / January / high-vol regimes, discovered after the pooled test failed. Sometimes real; always suspect when unplanned. 7. **HARKing** — Hypothesizing After Results are Known: the mechanism story written last, presented first. The tell is a theory section that predicts *exactly* the estimated coefficients and nothing else. ## The numerical intuition The flexibility arithmetic is unforgiving. Suppose each practice offers just 3 defensible variants and you exploit four of the seven: 3⁴ = 81 implicit specifications. At a 5% false-positive rate per spec, the probability *at least one* clears significance exceeds 98% under a true null — and that's before any [parameter grid](/guides/backtest-overfitting/#how-it-happens-in-practice-ranked-by-frequency) multiplies it further. Flexibility compounds like leverage. ## Tells a reader can spot in the published paper - **Fragile-boundary results**: significance that lives at p = 0.049, or an effect present at the 30% trim but not 25%. - **Asymmetric robustness sections**: ten checks that all pass, none that bind — real robustness sections contain [results that fail](/guides/robustness-trading-strategy/), because honest grids have edges. - **Precision without preregistration**: "we filter stocks below $4.83" — oddly specific thresholds are usually optimized ones. - **The moving outcome**: abstract claims returns; tables deliver risk-adjusted-alpha-on-a-subsample. - **Distribution of reported t-stats**: across the literature, t-values bunch suspiciously just above significance thresholds — a field-level fingerprint documented by the [replication studies](/guides/interpret-factor-zoo-paper/). ## The fixes, by role **As a researcher:** pre-register the specification in your [experiment log](/guides/quant-research-pipeline/#1-idea-intake--the-experiment-log) before touching data; treat every post-hoc change as a new trial in your [DSR count](/guides/deflated-sharpe-ratio/); report the full grid, including failures; and adopt the [evaluate-once holdout](/guides/walk-forward-out-of-sample-testing/#the-evaluate-once-rule) as constitutional law. **As a reader:** apply the t ≈ 3 hurdle for anything from a searched space; weight [post-publication performance](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer) over in-sample tables; and prefer papers whose robustness sections have casualties. ## Questions to ask when reading 1. Which choices in this paper *could have been made after seeing results* — and does anything indicate they weren't? 2. Is the headline specification the one a naive researcher would have run first? If not, why not? 3. What does the paper's own appendix reveal about the size of the searched space? 4. Do any reported checks actually threaten the result? ## What our scoring captures — and misses Our [rigor axis](/score-guide/) partially proxies for p-hacking resistance: papers scoring high tend to have disclosed grids, symmetric robustness sections, and out-of-sample discipline — the practices that make hacking hard. But p-hacking's essence is *invisible flexibility*, and no text-level score sees the specifications that were run and discarded. The score tells you the paper's visible process is sound; the seven-practices lens is how you price the invisible one. --- **The mechanics:** [Backtest overfitting →](/guides/backtest-overfitting/) · **The arithmetic:** [statistical vs economic significance →](/guides/statistical-vs-economic-significance/) · **The field-level view:** [factor-zoo papers →](/guides/interpret-factor-zoo-paper/) ---------------------------------------------------------------------- # Pairs Trading: Assumptions, Cointegration, and Failure Modes URL: https://thequant.space/guides/pairs-trading-failure-modes/ Section: Guides Date: 2026-09-07 Description: The assumptions under pairs trading and how each fails: cointegration's fragility, the selection problem, divergence risk, and the post-2000s decay of the classic approach. **Not an implementation claim.** The classic distance-method pairs strategy [famously decayed](/guides/map-statistical-arbitrage/) after its academic documentation; this guide teaches the mechanics and failure modes — the things that determine whether any *modern* variant deserves attention. **Definition.** Pairs trading: find two instruments whose prices move together, trade the spread when it diverges, profit on reconvergence. It's [stat-arb's](/guides/statistical-arbitrage-signal-to-portfolio/) minimal unit, and its clean structure makes every failure mode legible — which is why it remains the field's best teaching example even as its classic form retired. ## The three assumptions, and what tests them **1. The relationship is real.** Two testing traditions: **distance/correlation** (simple, finds co-movers, can't distinguish "related" from "were both trending") and **cointegration** (Engle-Granger, Johansen — tests for a *stationary spread*, the actual tradeable object). The trap in both: **selection across a large universe manufactures relationships**. Test 500 stocks pairwise and 5% of ~125,000 pairs pass at 5% significance by construction — [the multiple-testing problem](/guides/statistical-vs-economic-significance/#why-the-usual-significance-bar-is-broken-in-finance) wearing its most classic costume. Any pairs paper not correcting for search breadth across pairs is reporting [luck's order statistics](/guides/deflated-sharpe-ratio/). **2. It persists.** Cointegration measured in-sample is a *historical* statement; the economic linkage behind it (same industry, dual listing, holding structure) is what persists or doesn't. Pairs with a **mechanical linkage** (share classes, ADRs, index-vs-constituents) have persistent relationships and tiny, crowded spreads; pairs with only **statistical linkage** have wide spreads and relationships that [break with regimes](/guides/regime-dependence/). That inverse relation between persistence and profitability *is* the strategy's central trade-off, and papers silent on which type they trade are averaging over it. **3. Reversion arrives before capital leaves.** A stationary spread reverts *eventually*; your margin, borrow, and patience operate on deadlines. The half-life estimate (from an OU/error-correction fit) is the load-bearing number — and it's estimated with [wide error bars](/guides/confidence-intervals-strategy-research/) from limited data. A "20-day half-life" pair that draws a 6-month divergence isn't misbehaving statistically; you just met the distribution's tail with leverage on. ## The worked micro-example Spread = A − β·B, β from cointegrating regression. Enter at ±2σ, exit at 0. Three quiet leaks even here: **β is estimated** (its error means your "hedged" position carries directional exposure that [compounds with the divergence itself](/guides/portfolio-optimization-estimation-error/)); **σ is regime-dependent** (a 2σ entry in calm vol is a 0.8σ entry in crisis vol — entries cluster exactly when spreads are structurally widening); and **the exit at 0 assumes β still holds** at exit time. Rolling re-estimation helps and adds its own [turnover](/guides/turnover-factor-returns/) and whipsaw. ## Common ways this fails in production - **The false pair**: passed cointegration by search; diverges permanently on the first idiosyncratic event. Post-mortem always shows there was never an economic linkage. - **The broken pair**: real linkage, then M&A, index deletion, spin-off, or delisting severs it — [corporate actions](/guides/corporate-actions-adjusted-prices/) are the pair-killer papers rarely model, and the short leg's version is worst. - **The crowded unwind**: statistical pairs share membership across every quant book; a deleveraging event widens *all* spreads simultaneously — the [correlation-when-it-matters failure](/guides/high-sharpe-ratio-not-investable/#5-correlation-arrives-exactly-when-it-matters). - **Short-leg mechanics**: borrow recalls and buy-ins force-close the profitable side of the trade at the worst time. - **The 2σ trap in vol terms**: entry thresholds in *fixed* σ units auto-concentrate entries in high-vol regimes unless vol-scaled. ## Questions to ask a pairs paper 1. How many pairs were searched, and what correction was applied? 2. Mechanical or statistical linkage — and does the P&L come from the type claimed? 3. What happens to the equity curve if the worst 3 divergences run twice as long? 4. Is the sample [pre- or post-decay era](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer), and does the paper engage with the decay literature at all? ## What our scoring captures Cointegration papers split sharply on our axes: methodology papers (new tests, new spread models) sit in [Lab Rats](/score-guide/), and the rigor axis rewards exactly what this genre skips — search-breadth honesty, divergence-tail accounting, and cost-inclusive results. The [stat-arb map's](/guides/map-statistical-arbitrage/) pairs branch shows the ranking. --- **The general machinery:** [signal to portfolio →](/guides/statistical-arbitrage-signal-to-portfolio/) · **The statistics:** [multiple testing →](/guides/statistical-vs-economic-significance/) · [regime breaks →](/guides/regime-dependence/) ---------------------------------------------------------------------- # Point-in-Time Data Architecture for Backtesting URL: https://thequant.space/guides/point-in-time-data-architecture/ Section: Guides Date: 2026-09-07 Description: Designing a point-in-time data layer: bitemporal modeling, the as-of query pattern, a practical schema for prices, fundamentals, universes and events, and the migration path from a naive store. [Look-ahead bias](/guides/look-ahead-bias-point-in-time-data/) is usually treated as a backtest-code problem. It's a *schema* problem: a data layer that stores only current values **cannot** answer "what did we know on date t?", so every consumer reinvents (and fumbles) the correction. Build point-in-time into the architecture once and the whole [pipeline](/guides/quant-research-pipeline/#2-data--one-clean-layer-biases-handled-once) inherits correctness. This is the design guide. ## The core pattern: two timestamps, never one Every fact gets **event time** (when the fact is *about*: the fiscal quarter, the trade date) and **knowledge time** (when you *could have known it*: announcement, feed arrival, revision receipt) — the bitemporal pattern. Updates never overwrite; they append a new row with a later knowledge time. The invariant that makes backtests honest: **a query as-of knowledge-time K sees exactly the rows with knowledge_time ≤ K** — restatements, corrections, and backfills become *visible history* instead of silent rewrites. ```sql -- fundamentals, bitemporal CREATE TABLE fundamentals ( entity_id text, -- survivorship-safe id, not ticker (tickers get reused) period date, -- event time: fiscal period end field text, -- 'eps', 'revenue', ... value numeric, knowledge_ts timestamptz, -- when this value became knowable source text, PRIMARY KEY (entity_id, period, field, knowledge_ts) ); -- the as-of read every backtest uses SELECT DISTINCT ON (entity_id, period, field) * FROM fundamentals WHERE knowledge_ts <= :backtest_date ORDER BY entity_id, period, field, knowledge_ts DESC; ``` ## Schema by data class — where the pattern bites **Prices**: mostly single-temporal (the close is knowable at the close) — but store **unadjusted prices + a corporate-actions table**, because [adjustment factors are knowledge-timed facts](/guides/corporate-actions-adjusted-prices/): a split announced today changes the *adjusted* history, and pre-adjusted storage is a rolling rewrite of the past. Adjust at query time, as-of. **Fundamentals**: the full bitemporal treatment above — this is the class where [restatement leakage](/guides/look-ahead-bias-point-in-time-data/#1-restated-fundamentals--the-classic) lives. If your vendor sells only current-vintage values, `knowledge_ts = first_seen_at` from *your own ingestion* builds PIT forward from today — imperfect for history, airtight from now on. **Universe & membership**: an interval table (`entity_id, list, valid_from, valid_to, knowledge_ts`) — the [survivorship](/guides/survivorship-bias/#building-a-survivorship-clean-backtest) and [membership-leak](/guides/look-ahead-bias-point-in-time-data/#3-retroactive-universe-and-membership) killer. "Universe as of t" becomes a where-clause, not a research project. **Events**: [event studies](/guides/event-driven-research/#problem-1-defining-events-without-hindsight) need the event *record's* knowledge time (when the feed reported it) separate from the event's own timestamp — database backfill is a knowledge-time fact. Same table shape as fundamentals. **Reference data** (identifiers, listings, tick sizes): slowly-changing intervals, keyed by a permanent entity id. The id decision is load-bearing: **join everything on entity ids, map tickers per-date** — ticker reuse is a silent [wrong-join generator](/guides/reproduce-quant-research/#stage-2--align-match-their-world-before-testing-it). ## Making it usable: the loader contract PIT schemas die of inconvenience when every query needs the DISTINCT ON dance. The fix is the [shared loader](/guides/quant-research-pipeline/#2-data--one-clean-layer-biases-handled-once): one function — `load(fields, universe, as_of)` — that *only* exposes as-of reads, so leakage requires deliberately bypassing the API rather than forgetting a clause. Materialize common vintages (e.g., month-end snapshots) for speed; the bitemporal base stays the source of truth. Storage-wise this all lives happily in [Postgres](/guides/notebook-to-production/#2-postgres-is-the-source-of-truth-not-your-process-memory) at fundamentals scale, with the [tick layer](/guides/store-tick-data-efficiently/) staying in Parquet — the two-tier split from the [format guide](/guides/parquet-vs-database/). ## The migration path from a naive store 1. **Freeze and archive** the current-vintage store ([raw zone](/guides/quant-research-stack/) rules). 2. **Start capturing knowledge time now** — `first_seen_at` on every ingest; PIT correctness accrues forward immediately. 3. **Backfill where vintage sources exist** (vendor PIT products, your own historical pulls); label backfilled knowledge times as estimated. 4. **Move consumers to the as-of loader** one study at a time; the diff between old and new results is itself [a leakage measurement](/guides/look-ahead-bias-point-in-time-data/) worth writing down. ## Common ways PIT architectures fail - **PIT tables, non-PIT joins**: the fundamentals are bitemporal but the ticker-map or universe join isn't — one current-vintage join re-contaminates everything downstream. - **The bypass habit**: ad-hoc queries against base tables "just for exploration" leak into features; make the loader the path of least resistance. - **Vendor resets**: a full-history refresh from the vendor arrives all with today's knowledge time; distinguish *provider vintage* from *your receipt* or backfills poison the axis. - **Silent snapshot drift**: materialized vintages not rebuilt after late-arriving data disagree with the base — [reconcile them on schedule](/guides/notebook-to-production/#5-logs-metrics-and-the-alert-that-actually-matters). ## Questions to ask of any data layer (yours or a vendor's) 1. Can it answer "what did we know on t?" *by construction*, or by convention? 2. Are corporate actions facts-with-timestamps or baked-in adjustments? 3. What happens to a restatement — new row, or overwrite? 4. Do joins run on permanent ids? ## Where this connects This architecture is [the look-ahead guide](/guides/look-ahead-bias-point-in-time-data/) made structural, the prerequisite for [honest fundamental-factor research](/guides/evaluate-factor-investing-paper/), and the reason [vendor PIT capability](/guides/market-data-vendors/#2-is-fundamental-data-point-in-time) is worth paying for: vendors selling knowledge-timed data let you skip the hardest backfill. --- **The bias it kills:** [look-ahead & PIT →](/guides/look-ahead-bias-point-in-time-data/) · **The adjustment layer:** [corporate actions →](/guides/corporate-actions-adjusted-prices/) · **The storage substrate:** [Parquet vs Postgres →](/guides/parquet-vs-database/) ---------------------------------------------------------------------- # Portfolio Optimization: Estimation Error and Regularization URL: https://thequant.space/guides/portfolio-optimization-estimation-error/ Section: Guides Date: 2026-09-07 Description: Why mean-variance optimization amplifies estimation error, the error-maximization mechanism, the regularization toolkit (shrinkage, constraints, resampling), and the 1/N benchmark that keeps everyone honest. **Not an implementation claim.** This is the mechanics of why textbook optimization fails and what the working fixes cost — the context for reading the [portfolio-optimization literature](/topics/portfolio-optimization/), where "new optimizer beats mean-variance" is the genre's evergreen and [suspicious](/guides/evaluate-trading-backtest/) headline. **The mechanism.** Mean-variance optimization treats its inputs — expected returns μ and covariance Σ — as *known*. They're estimated, with error, and the optimizer is structurally an **error maximizer**: assets with overestimated returns or underestimated risk look attractive *because of the error*, so the optimizer concentrates there by design. Small input perturbations swing weights violently (the classic demonstrations flip 40-point allocations on 50 bps of μ), and out-of-sample, optimized portfolios routinely trail naive ones. ## The arithmetic of hopeless inputs Why μ is the problem: the [standard error of a mean return](/guides/confidence-intervals-strategy-research/#the-numbers-nobody-internalizes) shrinks with √T, and at portfolio-relevant precision (distinguishing 4% from 6% expected return) the required T runs to *decades per asset* — during which [the regime changed](/guides/regime-dependence/). Covariance is friendlier (more effective observations, especially at higher frequency) but degrades with dimension: estimating a 500×500 Σ from a few years of daily data leaves the sample eigenstructure part noise — and the optimizer [loads on the noisiest eigenvectors](/guides/robustness-trading-strategy/). Hence the field's practical hierarchy: **weights should depend heavily on Σ, barely on μ** — most working allocation schemes ([risk parity, min-variance](/guides/risk-parity-min-variance/)) are precisely μ-free constructions. ## The regularization toolkit - **Shrinkage on Σ** (Ledoit-Wolf lineage): blend sample covariance toward structure; the standard, nearly-free fix for dimension noise. - **Shrinkage on μ** (Black-Litterman lineage): anchor return estimates to an equilibrium prior, moving only where views justify — honest about μ's weakness rather than pretending it away. - **Constraints as implicit shrinkage**: long-only and position caps look crude and *act* statistical — binding constraints suppress exactly the extreme weights error creates. The practitioner's regularizer since before the theory named it. - **Resampling/bagging**: optimize over perturbed inputs, average the weights — smooths the error-chasing at computational cost. - **Turnover penalties**: error-chasing shows up as [rebalancing churn](/guides/turnover-factor-returns/#the-four-turnover-reduction-techniques-and-their-prices); penalizing it regularizes through time. Each buys stability by *biasing away from the optimizer's opinion* — the entire literature is a negotiation over how little to trust the inputs. ## The benchmark that keeps everyone honest: 1/N Equal weight uses no estimates, maximizes diversification-per-assumption, and famously survives comparison with a long list of optimized alternatives out of sample. It isn't magic — it's the zero-estimation-error corner solution, strongest exactly where estimation is weakest (broad equity universes) and weakest where structure is real and estimable (assets with order-of-magnitude vol differences, where [naive equal weight is an implicit risk bet](/guides/risk-parity-min-variance/)). Its role in evaluation is non-negotiable: **any optimization paper not benchmarking against 1/N and min-variance has skipped the controls** — the [dumb-baseline rule](/guides/evaluate-ml-trading-paper/#the-ml-paper-checklist) in allocation clothing. ## Common ways optimization fails in production - **The great-backtest optimizer**: in-sample optimized weights fit the sample's accidents; the [out-of-sample gap](/guides/walk-forward-out-of-sample-testing/) is the estimation error collected. - **Corner solutions on new data**: a fresh month moves estimates, the optimizer slams to new corners, and [turnover bills](/guides/turnover-factor-returns/) arrive for the noise. - **Crisis covariance**: correlations estimated in calm converge toward 1 in stress — the diversification the optimizer priced [disappears exactly when priced](/guides/high-sharpe-ratio-not-investable/#5-correlation-arrives-exactly-when-it-matters). Stress-conditional Σ, or at least stress-tested weights, is the fix papers skip. - **Constraint theater**: constraints tuned until the optimizer produces the portfolio you wanted anyway — at which point the optimization is [narrative, and its trials countable](/guides/backtest-overfitting/). ## Questions to ask an optimization paper 1. Is 1/N in the comparison table — and does the new method beat it *net of its own turnover*? 2. Are inputs estimated out-of-sample, walk-forward, with the weights' stability shown? 3. How does the method behave when fed *deliberately perturbed* inputs — [the sloppiness test](/guides/robustness-trading-strategy/#5-implementation-robustness--the-sloppiness-test) for allocators? 4. Does improvement survive at realistic rebalance frequencies? ## What our scoring captures The [portfolio hub](/topics/portfolio-optimization/) splits on exactly this guide's axis: derivation-heavy optimizer papers without out-of-sample horse races land in [Lab Rats](/score-guide/); the rigor tier belongs to papers with 1/N in the table, walk-forward inputs, and turnover-netted results. The score can't judge whether *your* universe has estimable structure — the 1/N-vs-structure question above — but it reliably surfaces the papers that tested theirs. --- **The μ-free family:** [risk parity & friends →](/guides/risk-parity-min-variance/) · **The error's bill:** [turnover →](/guides/turnover-factor-returns/) · **The input problem:** [confidence intervals →](/guides/confidence-intervals-strategy-research/) ---------------------------------------------------------------------- # Regime Dependence: When a Historical Edge Stops Working URL: https://thequant.space/guides/regime-dependence/ Section: Guides Date: 2026-09-07 Description: Regime dependence in trading strategies: why edges are conditional on market states, how to detect regime-carried backtests, decay vs regime-shift diagnosis, and what to do when an edge goes quiet. **Definition.** A strategy is regime-dependent when its edge is **conditional on a market state** — a volatility level, rate environment, trend phase, liquidity ecology, or participant mix — rather than unconditional. Almost every real edge is; the failure isn't dependence itself but *unexamined* dependence: a backtest that quietly averaged over a favorable mix of states, sold as an all-weather number. Regime dependence is also the honest competitor to the overfitting explanation: when a live edge goes quiet, [noise-fitting](/guides/backtest-overfitting/), regime shift, and ordinary variance are the three suspects, and they demand different responses. ## The worked example: the average that never happens A short-volatility carry strategy backtests at SR 1.6 over 2012–2019. Decompose by VIX tercile and the story changes: SR ≈ 2.5 in the calm two-thirds of months, SR ≈ −1.8 in the turbulent third. The headline 1.6 is an *average over a regime mix* — and the strategy's forward performance is entirely a bet on that mix persisting. 2020 reweights the mixture for one quarter and erases three years of carry. Nothing "stopped working"; the conditional structure was always there, and the unconditional Sharpe was [an answer to a question nobody would knowingly ask](/guides/high-sharpe-ratio-not-investable/#5-correlation-arrives-exactly-when-it-matters). The same decomposition indicts many trend, carry, and liquidity-provision results across [the archive](/topics/volatility/). ## Detecting regime-carried results in a paper - ☐ **Conditional performance table**: results by volatility tercile, rate environment, and trend state. Absence in a 15+ year backtest is the omission that [robustness sections owe](/guides/robustness-trading-strategy/#3-time-robustness--subperiods-and-regimes). - ☐ **Contribution concentration**: what fraction of total P&L comes from the best 10% of months? Edges earned in bursts are regime bets with a drumroll. - ☐ **Era honesty**: does the sample span structural breaks the mechanism cares about — decimalization for [microstructure](/guides/evaluate-microstructure-paper/#the-worked-example-true-here-false-there), the zero-rate decade for carry, pre/post-2018 crypto institutionalization? A mechanism whose habitat no longer exists has a perfect backtest and no future. - ☐ **Mechanism-regime coherence**: the paper's story should *predict* the conditional pattern. Momentum that claims a behavioral mechanism but earns only in one rate regime has a story-evidence mismatch. ## The live-trading diagnosis problem The hardest question in production: the strategy is down — **decayed, regime-shifted, or just unlucky?** They're separable only with pre-committed structure: - **Base rates first**: a true SR 1.0 strategy spends long stretches underwater with [substantial probability](/tools/equity-curve-simulator/) — run the fan before pattern-matching to doom. Most "it stopped working" calls at month 8 are [reading noise](/guides/confidence-intervals-strategy-research/#the-numbers-nobody-internalizes). - **Regime instruments**: if the edge is conditional, monitor the *conditioning variable*, not just P&L. Short-vol losing while VIX is at 35 is on-model; losing at VIX 12 is evidence against the mechanism. - **Decay signature**: alpha decay from crowding tends to be *gradual and monotonic* (each year weaker than the last, spreads compressing) versus regime losses which are *state-locked and reversible*. The [post-publication literature](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer) shows decay; 2020's factor reversals showed regime. - **Pre-commit the rules**: the [production checklist's](/guides/production-checklist/#gate-6--go-live-protocol) written stop criteria exist precisely because this diagnosis is impossible to make honestly mid-drawdown. ## What to do with a regime-dependent edge Dependence is manageable; three honest options, in order of ambition: **size by state** (condition exposure on the regime variable — now the regime model is part of the strategy and [counted in its trials](/guides/deflated-sharpe-ratio/)); **diversify across regime exposures** (pair edges with complementary habitats — the real argument for [multi-strategy construction](/topics/portfolio-optimization/)); or **accept and disclose** (run it unconditionally, but with the conditional table taped to the monitor and drawdown expectations set by the bad-state Sharpe, not the average). ## Questions to ask while reading 1. What state variable would this mechanism *predict* conditions the returns — and does the paper test it? 2. What share of profits comes from the single best regime episode? 3. Does the sample's regime mix resemble any plausible future? 4. If this edge went quiet for 18 months, what observation would distinguish shift from decay? ## What our scoring captures — and misses Subperiod and conditional analysis feed the [rigor axis](/score-guide/) directly — papers with regime tables outrank equal-Sharpe papers without them. What scoring can't do is judge *habitat survival*: whether the states an edge needs still occur is a forward-looking market question, not a paper-quality one. The score certifies the fossil is real; whether the climate still supports the animal is your call. --- **Companions:** [Robustness's five dimensions →](/guides/robustness-trading-strategy/) · [Drawdown base rates →](/tools/equity-curve-simulator/) · [When to stop a live strategy →](/guides/production-checklist/) ---------------------------------------------------------------------- # Reproducible Quant Research with Docker URL: https://thequant.space/guides/docker-reproducible-research/ Section: Guides Date: 2026-09-07 Description: Using Docker for reproducible quant research: what containers do and don't solve, the research-image pattern, determinism beyond the environment, and when containers are overkill. "It works on my machine" has a container-shaped solution; "it produces the same Sharpe" does not. Docker pins one axis of the [reproducibility triple](/guides/versioning-datasets-backtests/) — the **environment** — completely and portably. This guide is the research-flavored pattern, and an honest accounting of what containers *don't* fix. ## What the container actually pins A research image freezes: OS and system libraries (the BLAS/LAPACK builds that make [numerical results differ across machines](/guides/auditing-a-monte-carlo/)), the language runtime, and — via a lockfile *inside* the build — exact package versions. The contract: any machine, any year, `docker run` → the same bits execute. That kills the classic failures: the dependency upgrade that [silently changed results](/guides/versioning-datasets-backtests/#common-versioning-failures), the colleague who can't build your stack, and the six-months-later rerun on a machine that no longer exists. The research-image pattern, minimally: ```dockerfile FROM python:3.12-slim@sha256: # digest, not tag — tags move COPY pyproject.toml uv.lock ./ RUN pip install uv && uv sync --frozen # lockfile is the law COPY src/ ./src # code last: layer caching # no data in the image — data mounts, versioned separately ``` Three rules doing the work: **base images by digest** (`:3.12-slim` is a moving target); **lockfile-frozen installs** (the image build must not resolve dependencies — resolution is the nondeterminism you're killing); and **data stays outside** ([mounted, with its own versioning](/guides/versioning-datasets-backtests/#versioning-data-snapshots--content-addresses)) — images with baked-in data conflate two version axes and bloat unmanageably. ## What Docker does not solve The list that separates container users from reproducibility havers: - **Seeds and algorithmic nondeterminism**: same image, unseeded RNG → different results; GPU kernels and multi-threaded reductions add [nondeterminism a container happily preserves](/guides/evaluate-ml-trading-paper/#the-ml-paper-checklist). Seed everything; pin thread counts (`OMP_NUM_THREADS` and friends in the image's env); know which ops are nondeterministic by design. - **Data versions**: the container reproduces the *computation*, on whatever data you mount — [manifest-hashed snapshots](/guides/versioning-datasets-backtests/) are a separate discipline. - **Config drift**: parameters live outside the image ([as they should](/guides/notebook-to-production/#gate-2--code)); the run row records them. - **Hardware-order effects**: float summation order differs across core counts; bit-identical across *machines* sometimes requires pinning parallelism too. Decide your tolerance ([stated, like a reconciliation](/guides/reproduce-quant-research/#stage-4--reconcile-match-statistics-before-results)) before declaring divergence a bug. ## The honest overkill line Containers earn their keep at: **handoff boundaries** (a [replication package](/guides/reproduce-quant-research/#stage-6--document-make-your-reproduction-reproducible) someone else must run), **production** ([where they're already the runtime](/guides/notebook-to-production/#1-one-vps-containers-as-the-unit-of-operation) — research/live parity is a real win), **many-machine work** ([cloud GPU rentals](/guides/local-vs-cloud-gpu/) where the image *is* the setup), and **long-horizon archival** of published results. For a solo researcher iterating locally, **a lockfile + pinned interpreter delivers most of the value at none of the friction** — the honest default, with the Dockerfile added the day any boundary above appears. What's *not* optional at any scale is the lockfile itself; "pip install pandas" in a README is [an unversioned experiment](/guides/versioning-datasets-backtests/#versioning-backtests-the-run-record). ## Common containerized-research failures - **The tagged base image**: `FROM python:3.12` rebuilt a year later resolves different digests — reproducible-looking, wasn't. - **Resolution inside the build**: no lockfile → every build a new environment wearing the same name. - **The data-in-image antipattern**: 40GB images, version soup, and [backup policies](/guides/notebook-to-production/#6-backups-you-have-restored-at-least-once) that can't tell code from data. - **Container in research, bare metal in production**: the parity win [inverted](/guides/production-checklist/#gate-2--code) — two environments, one strategy, divergence by construction. - **Unpinned threads on shared runners**: CI produces different numerics than the workstation; the "bug hunt" is a core-count diff. ## Questions to ask 1. Does a from-scratch build today produce the same environment as six months ago? (Digest + lockfile = yes.) 2. Same image, same data, same seed, twice: identical outputs? If not, [what's the stated tolerance](/guides/auditing-a-monte-carlo/#61-test-invariants-not-outputs)? 3. Is the image digest in the [run row](/guides/versioning-datasets-backtests/#versioning-backtests-the-run-record)? ## What our scoring captures Environment specification is a [rigor](/score-guide/) micro-signal with macro correlation: papers shipping Dockerfiles or lockfiles nearly always clear the bigger reproducibility bars too — it's the visible tip of the [discipline stack](/guides/reproduce-quant-research/). In our archive the pattern concentrates in the [ML hub's](/topics/machine-learning/) top tier, where reruns are cheapest to demand and rarest to survive. --- **The other axes:** [versioning data & runs →](/guides/versioning-datasets-backtests/) · **Where images run:** [production stack →](/guides/notebook-to-production/) · [cloud GPUs →](/guides/local-vs-cloud-gpu/) ---------------------------------------------------------------------- # Research-to-Production Checklist for an Automated Trading Strategy (2026) URL: https://thequant.space/guides/production-checklist/ Section: Guides Date: 2026-09-06 Description: The complete checklist for taking a backtested strategy live: validation gates, execution safety, kill switches, monitoring, reconciliation, and the go-live protocol. Strategies rarely die from bad math. They die from a timezone bug that shifted signals by a day, an order loop that retried itself into triple position, a feed that went stale silently for a week. The gap between backtest and survival is operational, and operations respond well to checklists. This is ours — six gates, in order. Treat every unchecked box as a reason not to size up. ## Gate 1 — Validation (before writing any production code) - ☐ **Untouched out-of-sample period** evaluated exactly once, after all iteration stopped. If you peeked and iterated, it's training data now — carve a new one. - ☐ **Cost-sensitivity curve**: P&L at 1×, 2×, 3× your assumed costs. A strategy that dies at 2× costs is a cost model, not a strategy. - ☐ **Execution-lag sensitivity**: signals delayed by one bar/one minute/one hour. Fragility here predicts slippage pain later. - ☐ **Capacity estimate**: at what size does your fill assumption break? (Thin books make this the first constraint — see the [prediction-market guide](/guides/prediction-market-backtesting/).) - ☐ **Regime slices**: performance in at least three distinct volatility/trend regimes, and you can articulate *why* it should work in the ones where it does. - ☐ **Sample-size honesty**: trades clustered by underlying event/period counted as clusters, not independent observations. ## Gate 2 — Code - ☐ **One signal implementation.** The backtest and live system import the *same* signal function from the same module. Reimplementing "the same logic" for production is the most reliable bug generator in this business. - ☐ **Config out of code**: instruments, sizes, thresholds, endpoints in versioned config files; secrets in environment/secret storage, never in the repo (and never in shell history — `history | grep -i key` on your own machine is a sobering audit). - ☐ **Pinned, reproducible environment** (lockfile, container, or both). "Works on my machine" must be a provable statement. - ☐ **Idempotent startup**: the process can crash and restart at any moment and reconstruct its state from the broker/exchange, not from memory or local files it hopes are current. - ☐ **Clock discipline**: NTP-synced host, all internal timestamps UTC, timezone conversions only at display boundaries. Off-by-one-timezone is the classic silent killer of daily strategies. ## Gate 3 — Infrastructure - ☐ **A boring VPS** in a reliable region (or your broker's supported colocation if latency genuinely matters — for daily/hourly strategies it doesn't). Residential connections and laptops are for development only. - ☐ **Process supervision**: systemd unit or container with restart policy, memory limits, and logs shipped somewhere the process can't corrupt. - ☐ **Independent monitoring host** (can be a free-tier cron elsewhere): the machine watching the strategy must not be the machine running it. - ☐ **Deploy runbook**: documented, tested steps to deploy, roll back, and cold-start from nothing. If recovery lives in your head, you don't have recovery — you have a plan to have a bad week. ## Gate 4 — Execution safety (the account-savers) - ☐ **Kill switch, physically separate**: an independent process (ideally a different host) that flattens all positions and cancels all orders when triggered — manually, by heartbeat loss, or by breaching a hard drawdown/exposure line. It must not share code or state with the strategy it polices. This is the single highest-value item on this page. - ☐ **Order idempotency**: client order IDs on every order; on any uncertainty (timeout, disconnect), *reconcile before re-sending*. Retry loops without reconciliation are how bots buy the same position five times. - ☐ **Pre-trade limit checks**, enforced in code, per order and per day: max order size, max position, max notional, max order rate. Fat fingers happen in config files too. - ☐ **Partial-fill and reject handling** actually tested — by simulating them, not by hoping. Include exchange-halt and symbol-suspension paths. - ☐ **Rate-limit budgets** for every API with backoff; hitting a broker rate limit mid-liquidation is a bad time to learn about 429s. ## Gate 5 — Monitoring and reconciliation - ☐ **Heartbeats with escalation**: strategy emits liveness; the independent monitor alerts (phone-level, not email-level) on silence. Silence must be indistinguishable from emergency — a monitor that only reports success reports nothing. - ☐ **Data-staleness detection**: alert when feed timestamps stop advancing. A strategy trading yesterday's prices believes everything is fine. - ☐ **Daily reconciliation, automated**: your system's positions/P&L/fees vs. the broker's statement, every day, diff alerted. This catches execution bugs, fee surprises, and corporate-action mishandling within 24 hours instead of at tax time. Second-highest-value item here. - ☐ **Decision logging**: every signal, order, fill, and skip with inputs attached, retained. When live diverges from backtest — it will — this log is the only way to learn *why*. - ☐ **Live-vs-backtest tracking**: realized fills vs. modeled fills, realized signal P&L vs. simulated, reviewed weekly. Divergence trend is your earliest decay indicator. ## Gate 6 — Go-live protocol - ☐ **Paper period** (2–8 weeks) with the real execution path against live markets; gate on fill-model accuracy, not just P&L sign. - ☐ **Minimum-size live period**: real money, embarrassingly small, long enough to see every order state (partial, reject, cancel, halt) at least once. - ☐ **Written scale schedule with stop criteria**: size steps, the metric gates for each step, and the drawdown/divergence conditions that force a step *down*. Decide them now, while you're calm. - ☐ **Incident log from day one**: every anomaly, cause, fix. Three entries about the same component is a redesign signal. - ☐ **Quarterly kill-switch drill**: actually trigger it. An untested kill switch is a decoration. ## The meta-rule Every item above earns its place by the same logic our [paper scoring](/score-guide/) rewards empirical rigor: assume the system is lying to you until an independent mechanism confirms it isn't. Backtests overfit; processes crash; feeds stall; brokers disagree. Production readiness is the sum of independent checks that catch each lie early. *This checklist says what must be true; the how — the actual minimal VPS/Docker/Postgres stack with working configs — is in [From notebook to production](/guides/notebook-to-production/). This checklist is general; regulated entities have compliance layers beyond its scope.* --- **Take it with you:** download the plain-markdown checklist — drop it into your repo or issue tracker. The newsletter below gets you the weekly research digest and future checklist revisions. ---------------------------------------------------------------------- # Risk Parity, Minimum Variance, and Maximum Diversification URL: https://thequant.space/guides/risk-parity-min-variance/ Section: Guides Date: 2026-09-07 Description: The μ-free allocation family compared: what risk parity, minimum variance, and maximum diversification each optimize, the leverage that makes risk parity work, and the shared failure modes. **Not an implementation claim.** These are the working members of the μ-free family from [the estimation-error guide](/guides/portfolio-optimization-estimation-error/) — allocation without return forecasts. This guide is what each *actually* optimizes and bets on, which their marketing rarely states. ## The three, precisely **Minimum variance**: minimize w'Σw. Pure risk minimization — concentrates in low-vol, low-correlation assets, which in equity universes means the low-volatility corner (and its [factor exposure](/guides/alpha-beta-alternative-risk-premia/): min-var equity portfolios are, in large part, the low-vol premium in optimizer clothing). Estimation burden: Σ only, and mostly its top structure — the family's most robust member. **Risk parity**: equalize *risk contributions* — each asset (or asset class) contributes equally to portfolio variance, so low-vol assets get large weights. The inseparable second half: unlevered risk parity across stocks/bonds is mostly bonds and underperforms by construction, so **risk parity is a leverage strategy** — lever the balanced-risk portfolio to a target vol. Its historical record therefore embeds a *financing-cost and bond-regime bet*: decades of falling rates flattered it; [the 2022 regime](/guides/regime-dependence/#detecting-regime-carried-results-in-a-paper) — stocks and bonds down together, financing costs rising — was its documented stress case, right on schedule. **Maximum diversification**: maximize the diversification ratio (weighted average vol ÷ portfolio vol) — tilting toward assets whose *correlation* structure adds the most independent risk. Sits between min-var and equal-weight; its bet is that the correlation matrix's structure is real and stable, [the exact quantity that degrades in stress](/guides/portfolio-optimization-estimation-error/#common-ways-optimization-fails-in-production). The unifying view: all three are answers to "what should weights depend on, if not μ?" — variance level (min-var), risk budgets (parity), or correlation structure (max-div). Each is a *different implicit bet* dressed as neutrality, and choosing between them is choosing which estimate you trust. ## The comparison that matters | | Min variance | Risk parity | Max diversification | 1/N | |---|---|---|---|---| | Estimates needed | Σ (top structure) | vols + a leverage rule | full correlation matrix | none | | Implicit factor bet | Low-vol premium | Bond duration + financing | Correlation stability | Small-cap/equal tilt | | Leverage | No | **Essential** | Optional | No | | Stress behavior | Best of family | Financing + correlation squeeze | Correlation convergence hurts most | Transparent | | [Turnover](/guides/turnover-factor-returns/) | Moderate | Vol-tracking driven | Highest (correlation-chasing) | Lowest | ## Common ways the family fails in production - **The 2022 mode** (risk parity): the diversifying asset and the growth asset fall together while leverage costs rise — three assumptions failing jointly, [as regime shifts do](/guides/regime-dependence/). - **Vol-tracking churn**: parity and min-var re-estimate risk continuously; vol spikes force deleveraging into falling markets — the family shares [volatility targeting's](/guides/volatility-targeting/) procyclical sell-low mechanics. - **Correlation-matrix fragility** (max-div especially): weights chase estimated correlation structure that is [part noise at realistic sample sizes](/guides/portfolio-optimization-estimation-error/#the-arithmetic-of-hopeless-inputs); stress converges correlations and deletes the claimed diversification. - **Leverage plumbing** (parity): financing spreads widen and margin terms tighten exactly in the drawdowns — [the urgency-convexity of costs](/guides/high-sharpe-ratio-not-investable/#6-costs-and-slippage-are-convex-in-urgency) applied to the whole portfolio. - **Backtest-era dependence**: the family's flagship long-run results were earned in a forty-year bond bull; [subperiod tables](/guides/robustness-trading-strategy/#3-time-robustness--subperiods-and-regimes) are not optional here. ## Questions to ask a paper in this family 1. Which estimates does the scheme actually consume, and how were they [computed out-of-sample](/guides/walk-forward-out-of-sample-testing/)? 2. For parity: what financing cost is assumed, and what happens at +200 bps? 3. Is 1/N (and min-var, for the fancier schemes) [in the comparison](/guides/portfolio-optimization-estimation-error/#the-benchmark-that-keeps-everyone-honest-1n)? 4. How did the scheme behave in stocks-bonds-correlated stress — in and out of sample? 5. What's the turnover, and [who pays it](/guides/turnover-factor-returns/)? ## What our scoring captures The [portfolio hub's](/topics/portfolio-optimization/) risk-based-allocation papers score well on rigor when they do the three honest things: subperiod (regime) tables, financing-cost accounting for levered schemes, and naive benchmarks. The recurring low-rigor pattern is the forty-year gross backtest ending before the regime that tests the scheme's core bet — which the [scoring axis](/score-guide/) penalizes as incomplete evidence, and this guide lets you name precisely. --- **The estimation problem underneath:** [optimization & estimation error →](/guides/portfolio-optimization-estimation-error/) · **The leverage layer:** [volatility targeting →](/guides/volatility-targeting/) ---------------------------------------------------------------------- # Statistical Arbitrage: From Signal to Executable Portfolio URL: https://thequant.space/guides/statistical-arbitrage-signal-to-portfolio/ Section: Guides Date: 2026-09-07 Description: The mechanics between a stat-arb signal and an executable portfolio: neutralization, sizing, netting, turnover control, and the places the translation destroys the edge. **Not an implementation claim.** This guide teaches the *mechanics* that determine whether a stat-arb idea can become a portfolio — not a strategy that makes money. The mechanics are the filter: most published signals die in the translation described below. A [statistical-arbitrage](/topics/statistical-arbitrage/) paper typically ends where the hard part begins: "we form a long-short portfolio from the signal." That sentence hides five engineering decisions, each with more P&L consequence than most signal refinements. ## Step 1: Neutralization — deciding what you're betting on A raw signal correlates with everything: market, sectors, size, volatility. Neutralize by regression or constraint until what remains is the bet you *intend*. The trade-offs are real: every neutralization costs gross exposure and adds turnover, and over-neutralizing (against 50 factors) can leave nothing but noise plus costs. The test from [the alpha-beta framework](/guides/alpha-beta-alternative-risk-premia/): regress your portfolio's returns on standard factors — the residual is what you built; if it's small, the signal was [beta wearing makeup](/guides/evaluate-trading-backtest/). ## Step 2: Sizing — signal strength is not position size Three standing schemes: **rank-based** (buckets — robust, throws away magnitude), **z-scored proportional** (uses magnitude — sensitive to outliers exactly when outliers are dangerous), and **optimizer-mediated** (feeds signal into [mean-variance machinery](/guides/portfolio-optimization-estimation-error/) — maximal theory, maximal [estimation-error leverage](/guides/portfolio-optimization-estimation-error/)). The stat-arb-specific trap: signals are strongest in the least tradeable names — small caps, wide [spreads](/guides/transaction-costs-slippage-market-impact/), scarce borrow — so *unconstrained* signal-proportional sizing concentrates the book precisely where [capacity](/guides/capacity-constraints/) is worst. Liquidity-scaled sizing (cap positions at a fraction of ADV) is the boring fix that keeps backtests honest. ## Step 3: Netting and the portfolio's real bet Cross-sectional books net: your 200 longs and 200 shorts share factor exposure, and the *net* book is much smaller than gross. Consequences papers skip: **margin and financing** are charged on gross while returns accrue on net — the leverage-to-edge ratio determines whether financing eats the alpha; **borrow** on the short side is a [real, name-specific cost](/guides/evaluate-factor-investing-paper/#the-factor-paper-checklist) — hard-to-borrow names cluster in the short decile of most signals; and **crowding correlation**: your neutralized book correlates with every other stat-arb book built from similar signals, which is invisible daily and decisive during [deleveraging cascades](/guides/regime-dependence/) (August 2007 remains the canonical exhibit). ## Step 4: Turnover control — the edge's tax rate Signal updates want trades; costs bill them. The mechanics that reconcile: **rebalance bands** (trade only when drift exceeds a threshold — cuts turnover 30–60% with minor tracking cost), **signal smoothing** (averaging fast signals before sizing), and **trade netting across signals** (one book from many signals nets internal crossings for free). The [turnover arithmetic](/guides/turnover-factor-returns/) rules: annual cost = turnover × per-side cost — engineer the left factor, since the right one is [the market's to set](/guides/transaction-costs-slippage-market-impact/). ## Common ways this fails in production - **The convergence assumption meets a divergence regime**: spreads that "always" mean-revert widen through your stop during risk-off; the loss arrives exactly when [borrow gets recalled and margin tightens](/guides/high-sharpe-ratio-not-investable/#5-correlation-arrives-exactly-when-it-matters). - **Execution-order asymmetry**: entering legs sequentially leaves you naked directional risk between fills; a "market-neutral" book is only neutral *after* both sides fill. - **Universe drift**: the live universe (borrowable, liquid, listed) shrinks relative to the backtest's, concentrating the book and amplifying single-name events. - **Signal decay masquerading as variance**: monitor the *cross-sectional IC*, not just P&L — [decay is monotonic; regimes are state-locked](/guides/regime-dependence/#the-live-trading-diagnosis-problem). - **Corporate actions in the shorts**: buyouts announced on your short leg are step-function losses no spread model contains. ## Questions to ask a stat-arb paper 1. Is the reported return *after* the neutralizations its own risk table implies? 2. What does the portfolio's short leg look like as a borrow list? 3. What's the signal's half-life vs the turnover its capture requires? 4. Does performance survive liquidity-capped sizing? ## What our scoring captures Papers that carry their signal all the way to a costed, constrained portfolio score materially higher on [rigor](/score-guide/) — the translation steps are exactly what "backtest-ready" means in our rubric. The [map of the field](/guides/map-statistical-arbitrage/) shows how few papers reach step 4; those that do cluster at the top of the hub. --- **The narrower classic:** [pairs trading's failure modes →](/guides/pairs-trading-failure-modes/) · **The tax:** [turnover →](/guides/turnover-factor-returns/) · **The ceiling:** [capacity →](/guides/capacity-constraints/) ---------------------------------------------------------------------- # Statistical Significance vs Economic Significance in Trading Research URL: https://thequant.space/guides/statistical-vs-economic-significance/ Section: Guides Date: 2026-09-07 Description: Why t-statistics mislead in finance: multiple testing and the factor zoo, the t > 3 hurdle, economic magnitude after costs, and the two-by-two matrix for judging any empirical result. Two findings: Strategy A earns 0.4 bps per trade with t = 4.2 on two million observations. Strategy B earns 40 bps per trade with t = 1.8 on three hundred. Which is real? Both. Which matters? Neither question is answered by the t-stat — A's edge disappears into [half a spread](/guides/transaction-costs-slippage-market-impact/), and B needs more data. Statistical and economic significance are different questions, and the field's chronic error is answering the first while claiming the second. ## Why the usual significance bar is broken in finance The textbook t > 2 threshold assumes you ran **one** test. Nobody runs one test. Academia has published hundreds of return "factors"; each represents many unreported variations tried along the way. Test enough random signals and 5% clear t > 2 *by construction* — the **factor zoo** problem. The influential response (Harvey, Liu & Zhu's cross-section work) concluded that, accounting for the field's collective data mining, a new factor should clear **t ≈ 3.0 or higher** — and large-scale replication studies applying consistent methodology found that a substantial fraction of published anomalies fail even at conventional thresholds. The individual-researcher version is worse, because your trial count is invisible: fifty parameter combinations tried is fifty tests, and your best backtest is the *maximum* of fifty draws. What the maximum of N draws looks like by luck alone has a formula — that's the [deflated Sharpe ratio](/guides/deflated-sharpe-ratio/), this guide's quantitative sequel. The honest practice it implies: **count and report your trials.** A t of 2.5 from three pre-registered variants outranks a t of 3.5 from a 2,000-run grid search. Compounding it: financial data violates the assumptions under the t-stat — overlapping observations, [correlated bets counted as independent](/guides/evaluate-trading-backtest/#4-how-many-independent-bets-is-this-really), fat tails, and regime dependence all inflate naive significance. Clustered errors and block bootstraps help; they rarely rescue a marginal result. ## Economic significance: the four questions t-stats can't answer 1. **Magnitude after costs.** Take the effect size, subtract [realistic costs for the asset class](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges), multiply by achievable exposure. Many famous anomalies are statistically robust and economically extinct — real patterns, too small to trade. 2. **Capacity.** An edge confined to micro-caps or thin books can be simultaneously true and irrelevant beyond $2M — the [square-root law](/guides/transaction-costs-slippage-market-impact/#the-four-components-separated) guarantees it. 3. **Accessibility.** Does capturing it require shorting hard-to-borrow names, trading at the close auction at scale, or data nobody can buy? A true effect you cannot access prices at zero. 4. **Decay.** Published effects shrink — the well-documented post-publication decay in [factor research](/topics/factor-investing/) — so the relevant magnitude is the *post-publication, post-cost* one, which is also why reproducing older papers on fresh data is [the cleanest out-of-sample test available](/guides/reproduce-quant-research/#stage-5--stress-the-part-the-paper-didnt-do). ## The two-by-two that organizes every empirical claim | | **Economically large** | **Economically small** | |---|---|---| | **Statistically strong** | The real thing — rare; check [costs](/guides/transaction-costs-slippage-market-impact/), [biases](/guides/survivorship-bias/), capacity, then [reproduce it](/guides/reproduce-quant-research/) | "True but tiny" — academically interesting, a trap for practitioners; most surviving anomalies live here | | **Statistically weak** | "Promising but unproven" — the interesting quadrant; gather data, don't dismiss (small samples doom real effects too) | Noise — walk away regardless of narrative quality | The top-right trap deserves emphasis because it's where sophisticated readers get caught: a mountain of observations makes tiny effects gleam with enormous t-stats. High-frequency datasets are factories for this quadrant — two million ticks will find you a 0.3-bp effect at t = 6 that no fee schedule on earth lets you keep. ## How this maps to our scores The [empirical-rigor axis](/score-guide/) on every [paper page](/flowcharts/) is, in large part, this guide operationalized: papers rank higher for reporting effect sizes with costs, disclosing trial counts, testing subperiods, and distinguishing what's significant from what's tradeable. Papers with spectacular t-stats and silent turnover columns rank lower — now you know exactly why. --- **The quantitative sequel:** [The deflated Sharpe ratio →](/guides/deflated-sharpe-ratio/) · **The sample-size discipline:** [walk-forward & OOS testing →](/guides/walk-forward-out-of-sample-testing/) ---------------------------------------------------------------------- # Survivorship Bias in Quantitative Finance: How Dead Companies Fake Alpha URL: https://thequant.space/guides/survivorship-bias/ Section: Guides Date: 2026-09-07 Description: How survivorship bias inflates backtests, how large the effect is by asset class, the five-dead-tickers test for any dataset, and how to build survivorship-clean universes. Run a momentum backtest on the current S&P 500 membership back to 2000 and you will make money, because you have quietly required every company in the test to have *survived until today*. Enron, Lehman, and four hundred delisted small caps have been erased from history — along with the losses a real portfolio would have taken holding them. Survivorship bias is the most mechanical way a backtest lies, which makes it the first check in our [backtest evaluation framework](/guides/evaluate-trading-backtest/) and a fixture in the papers that score poorly on [empirical rigor](/score-guide/). ## The mechanism, precisely The bias enters wherever a historical universe is constructed from a *present-day* list: - **Delisted securities vanish** — and delisting is not random. Bankruptcies and failed listings exit with large negative final returns; using a survivors-only universe deletes exactly the left tail your strategy would have eaten. - **Index membership is applied retroactively** — "S&P 500 stocks since 2000" using *today's* membership tests only companies that grew into the index. Correct construction uses membership *as of each date*. - **Fund databases bury their dead** — hedge fund and mutual fund datasets historically dropped closed funds; reported category averages are averages of the funds that didn't blow up. - **Crypto has it worst**: dead tokens, delisted pairs, and defunct exchanges disappear from APIs entirely, taking their price histories with them ([our crypto hub](/topics/crypto-defi/) is full of backtests that never met a dead exchange). ## How big is it? Magnitudes worth carrying in your head (long-run US equity estimates; direction generalizes, size varies): - Long equity strategies on survivors-only data: roughly **1–4%/yr** of phantom return, worst in small caps where delisting rates are highest. - Strategies that *select for distress* — deep value, low price, high volatility — are distorted far more, because their picks are overrepresented among future delistings. A "buy losers" reversal backtest on survivors-only data is almost pure artifact. - Hedge-fund index returns: the classic estimates put survivorship (plus backfill) inflation at **2–4%/yr** — a large fraction of the headline "hedge fund premium." The asymmetry matters: the bias always flatters *long* exposure to the deleted tail, so it manufactures alpha for long strategies and can mask it for shorts. ## The five-dead-tickers test The fastest audit of any dataset, from our [market-data guide](/guides/market-data-vendors/#1-does-history-include-the-dead): pick five well-known corpses — a bankruptcy (Lehman), an acquisition (Twitter), a delisted small cap, a dead SPAC, a failed token — and query their full histories. Three grades: 1. **Full history including final trading days** → usable. 2. **History exists but ends months before the actual delisting** → partial credit; final-period returns (often the worst ones) are missing. 3. **Ticker not found** → the dataset is for dashboards, not research. Run it *before* subscribing to any vendor and before trusting any public dataset. It takes ten minutes and settles what marketing pages never state. ## Building a survivorship-clean backtest - **Universe as-of-date**: reconstruct membership from historical constituent files or point-in-time listing status — the same [PIT discipline](/guides/look-ahead-bias-point-in-time-data/) that governs fundamentals. - **Include delisting returns**: a position in a delisted stock must realize its final return — including the convention problem that final prices are sometimes missing (academic practice imputes substantial negative returns for performance-related delistings; whatever you choose, choose explicitly and disclose it). - **Source accordingly**: CRSP is the academic gold standard precisely because of delisting coverage; among practitioner vendors, delisting depth is the sharpest quality divider — [tiers and tests here](/guides/market-data-vendors/). - **Archive forward**: from today on, snapshot your own universe daily into your [raw zone](/guides/quant-research-stack/#layer-1-market-data--own-your-pipeline-rent-the-feed). Nobody can retroactively delete a survivor list you recorded yourself — the same "record your own books" logic as [prediction markets](/guides/prediction-market-backtesting/#2-phantom-liquidity). ## Reading papers with this lens When a paper doesn't state its survivorship handling, the default assumption is *contaminated* — authors who did the work say so, because it's expensive. Discount accordingly, and discount hardest for [factor-investing results](/topics/factor-investing/) in small caps and "loser rebound" effects, where the bias concentrates. This is one of the checks behind the rigor axis on every [paper page](/flowcharts/) we publish. --- **Next in the bias family:** [Look-ahead bias & point-in-time data →](/guides/look-ahead-bias-point-in-time-data/) · **Test your vendor:** [market-data guide →](/guides/market-data-vendors/) ---------------------------------------------------------------------- # The Deflated Sharpe Ratio, Explained with Real Numbers URL: https://thequant.space/guides/deflated-sharpe-ratio/ Section: Guides Date: 2026-09-07 Description: How the deflated Sharpe ratio corrects for multiple testing, fat tails, and short samples: the intuition, the formulas, and the worked example — 100 random backtests produce a Sharpe ≈ 2.5 by luck alone. Here is the number that should hang over every backtest review: **run 100 strategies with zero skill on one year of daily data, and the best of them will show an annualized Sharpe ratio of roughly 2.5 — by pure luck.** Not occasionally; *in expectation.* The deflated Sharpe ratio (DSR), from Bailey and López de Prado, is the tool that turns this fact into a hurdle: given how many things you tried, how good would your best result look if nothing worked? Beat that, or you've measured your own persistence. It is the quantitative sequel to [statistical vs economic significance](/guides/statistical-vs-economic-significance/): that guide says count your trials; this one says what the count costs you. ## Step 1: a measured Sharpe is a noisy estimate An estimated Sharpe ratio has sampling error like any statistic. For roughly i.i.d. returns, the standard error of an annualized Sharpe estimated from one year of daily data is **about 1.0** — a true-zero strategy will regularly print ±1 over a year. Three practical consequences before any multiple-testing correction: short track records say almost nothing; negative skew and fat tails (which trading strategies produce chronically — selling vol *manufactures* smooth-then-cliff return profiles) inflate measured Sharpes' reliability further; and the **probabilistic Sharpe ratio (PSR)** handles all of this by asking the right question — *what is the probability the true Sharpe exceeds some benchmark, given the sample length, skew, and kurtosis of my returns?* The PSR formula (per-period units): ```text PSR = Φ( (SR − SR*) · √(T−1) / √(1 − γ₃·SR + ((γ₄−1)/4)·SR²) ) ``` where T is sample length, γ₃ skewness, γ₄ kurtosis, and SR* the benchmark to beat. Negative skew and excess kurtosis shrink the denominator's information content — exactly the penalty smooth-until-they're-not strategies deserve. ## Step 2: the maximum of N tries is not a fair draw Now the multiple-testing half. If you evaluate N strategy variants and keep the best, you've sampled the *maximum* of N noisy Sharpes — and the expected maximum grows with N even when every true Sharpe is zero. Extreme-value theory gives the expected best-by-luck: ```text E[max SR] ≈ σ_SR · [ (1−γ)·Z(1 − 1/N) + γ·Z(1 − 1/(N·e)) ] ``` with γ ≈ 0.5772 (Euler–Mascheroni), Z the inverse normal CDF, and σ_SR the dispersion of Sharpe estimates across trials. The worked example behind this guide's opening number: with one year of daily data, σ_SR ≈ 1.0 annualized, and N = 100 gives E[max SR] ≈ 1.0 × [0.4228·Z(0.99) + 0.5772·Z(0.99632)] ≈ 1.0 × (0.98 + 1.55) ≈ **2.5**. The scaling is logarithmic — going from 100 to 1,000 trials only pushes the luck-ceiling toward ~3.25 — but almost nobody's true trial count is small: every parameter grid cell, feature set, universe tweak, and abandoned variant counts, and so do [respent holdouts](/guides/walk-forward-out-of-sample-testing/#the-evaluate-once-rule). ## Step 3: deflate The DSR is simply the PSR evaluated **against the luck-ceiling as the benchmark**: SR* = E[max SR] for your N. It answers, in one probability: *given my sample length, my return distribution's tails, and how many things I tried, how likely is it that my best backtest reflects real skill rather than selection?* A DSR near 1 survives; a spectacular raw Sharpe with a DSR near 0.5 is a coin flip wearing a tuxedo. ## Using it honestly, without ceremony - **Estimate N generously.** Include the informal tries. If you can't reconstruct it, use the grid you *would* have run — the estimate that flatters you is the wrong one. Correlated trials reduce the *effective* N (a formal treatment clusters trials by correlation); the honest shortcut is to count distinct strategy families at full weight. - **Log trials as you go.** An experiment log — one row per variant evaluated, part of the [research-pipeline discipline](/guides/quant-research-pipeline/) — makes N a database query instead of an argument with yourself. - **Apply it when reading, too.** A paper reporting one golden configuration from an unstated search space earns an N in the hundreds by default; this is priced into our [rigor scores](/score-guide/), and it's why single-configuration papers with no sensitivity analysis rank poorly across the [factor](/topics/factor-investing/) and [ML](/topics/machine-learning/) hubs. - **See it, once.** Ten minutes with our [equity-curve simulator](/tools/equity-curve-simulator/) — run 100 zero-edge paths and look at the best one — teaches the maximum-of-N intuition more permanently than any formula. ## The limits DSR is a hurdle, not a verdict: it corrects selection *within your recorded trials*, not [survivorship in your data](/guides/survivorship-bias/), [look-ahead in your pipeline](/guides/look-ahead-bias-point-in-time-data/), or [costs you didn't model](/guides/transaction-costs-slippage-market-impact/) — a leaked backtest deflates beautifully and is still fiction. It sits at the end of the evaluation chain, after [the eight backtest questions](/guides/evaluate-trading-backtest/), as the final filter on results that survived everything else. *References: Bailey & López de Prado, "The Deflated Sharpe Ratio" (2014); Harvey, Liu & Zhu on the cross-section of expected returns and the t ≈ 3 hurdle.* --- **The discipline that feeds it:** [trial counting →](/guides/statistical-vs-economic-significance/) · [evaluate-once holdouts →](/guides/walk-forward-out-of-sample-testing/) · **See luck's distribution:** [equity-curve simulator →](/tools/equity-curve-simulator/) · **Run your numbers:** [deflated Sharpe calculator →](/tools/deflated-sharpe-calculator/) ---------------------------------------------------------------------- # The Most Important Quant Finance Papers, by Topic URL: https://thequant.space/guides/important-quant-finance-papers/ Section: Guides Date: 2026-09-07 Description: The strongest quantitative finance papers in each research area, ranked by empirical rigor and math complexity from a 5,000+ paper scored archive — updated automatically as new work is scored. "Which papers should I read for X?" — this page is the archive's standing answer. Every section below lists the top-rated papers in one research area, ranked by our [rigor-weighted score](/score-guide/) (60% empirical rigor, 40% math complexity), and **regenerates automatically** as the [daily pipeline](/about/) scores new work — so unlike every static "essential papers" listicle, it doesn't fossilize. Two honest caveats. *Important ≠ famous*: this ranks evidence quality within our archive, not citation counts — the canonical textbook papers are covered by every syllabus already, while high-rigor recent work is exactly what's hard to find. And *rankings inherit the scoring system's limits*: the score can't see [unpublished trial counts](/guides/backtest-overfitting/) or [pipeline leakage](/guides/evaluate-ml-trading-paper/), so read winners with the [standard checklist](/guides/evaluate-trading-backtest/) anyway. ## Statistical arbitrage ## Market microstructure ## Machine learning for forecasting ## Reinforcement learning for trading ## Options & derivatives ## Volatility ## Portfolio optimization ## Factor investing & asset pricing ## Risk management ## High-frequency trading & execution ## Crypto & DeFi ## NLP & LLMs in finance ## Fixed income ## Stochastic control & optimal stopping ## Insurance & actuarial risk ## Commodities & energy ## Market simulation & agent-based models --- **Go deeper:** each area has a full [topic hub](/topics/) with every paper ranked, and several have [visual research maps](/guides/map-statistical-arbitrage/) showing the field's structure. **Reading order:** [how to read a quant paper →](/guides/how-to-read-quant-finance-papers/) · [what to replicate first →](/guides/choose-papers-to-replicate/) ---------------------------------------------------------------------- # TimescaleDB vs ClickHouse vs DuckDB vs kdb+ for Tick Data Research (2026) URL: https://thequant.space/guides/tick-data-databases/ Section: Guides Date: 2026-09-06 Description: An honest comparison of TimescaleDB, ClickHouse, DuckDB, QuestDB, and kdb+ for storing and querying tick data in quant research — by workload, not by benchmark marketing. Every database in this comparison can store a billion ticks. The differences that matter are what happens when you *query* them — and vendor benchmarks won't tell you, because each engine wins its own benchmark. What follows is a comparison by **workload shape**, based on each system's documented architecture and the patterns that repeat across quant teams. Where we make a performance claim, it's an architectural consequence, not a number from a marketing page; run your own benchmark on your own queries before committing (a starter harness is sketched at the end). ## Start with your workload, not the database Tick-data work stresses three very different capabilities: 1. **Research scans** — "average spread by symbol-hour across three years": full-table analytical scans, massively parallel, read-only. Column stores dominate. 2. **Live ingestion + recent-data queries** — streaming millions of events/day while querying the last hour: write throughput plus fresh-data visibility. 3. **Time-series joins** — as-of joins ("last quote at or before each trade"), window logic, order-book reconstruction: this is where general-purpose SQL engines historically embarrass themselves and specialized engines shine. Score each candidate against *your* mix; almost every wrong choice in this space is a category-2 tool bought for a category-1 job or vice versa. ## The candidates ### DuckDB — the research default In-process analytical engine over Parquet files; zero ops. For a single researcher doing historical analysis, it is simply the correct starting point in 2026: full SQL including a native **`ASOF JOIN`**, vectorized scans that saturate NVMe, and Arrow-native interchange with Polars/pandas so results move without serialization. Its honest limits: it's single-node, single-writer, and not an ingestion server — pair it with a raw Parquet archive (see the [stack guide](/guides/quant-research-stack/)) rather than treating it as a live store. Many teams discover their "we need a cluster" problem was actually a "we needed partitioned Parquet and DuckDB" problem. ### ClickHouse — the scale-up column store The open-source heavyweight for analytical scans over tens of billions of rows: aggressive compression, brutal parallelism, mature replication, and a huge operational knowledge base. It has an `ASOF JOIN`, materialized views for continuous aggregation, and handles high-rate inserts well in batches. Costs: it's a server you operate (or pay ClickHouse Cloud to), its SQL dialect has sharp edges, and mutation/update patterns are awkward by design — model your data append-only. Choose it when data or concurrency genuinely outgrows one process. ### TimescaleDB — Postgres with a time-series engine A Postgres extension adding hypertables, columnar compression, and continuous aggregates. Its killer feature is the ecosystem: full Postgres SQL, joins against your reference/positions tables, every ORM and BI tool, transactional updates. For *pure* tick-scan analytics it's architecturally behind the column stores (row-based hot storage, compression as a maturing add-on), and there's no native as-of join — you emulate with `LATERAL`, which works but won't win races. Choose it when tick data must live *next to* relational data and operational simplicity beats peak scan speed. (Licensing note: the compressed/columnar features live under the Timescale license, not vanilla Postgres open source — check terms if you redistribute.) ### QuestDB — ingestion-first time series Purpose-built for exactly this domain: very high streaming ingestion via InfluxDB line protocol, SQL with native **`ASOF JOIN`** and time-partitioned storage, minimal configuration. It's the strongest open-source answer to "I need to ingest live feeds all day *and* query them in SQL." The trade-offs are a younger ecosystem, thinner distributed story, and less battle-tested operational tooling than ClickHouse/Postgres. Well suited to a small desk capturing its own feed. ### kdb+/q — the incumbent Still the reference architecture on institutional desks: column store, memory-mapped historical partitions, and q — a language where as-of joins (`aj`) and time-series transforms are primitives, not workarounds. Nothing matches its expressiveness-per-line for order-book work. The costs are equally real: commercial licensing (list pricing is "talk to sales"; budget accordingly), a language with a steep learning curve and a thin hiring pool outside finance hubs, and an ecosystem that assumes you're an institution. For a solo shop it's rarely rational unless you already speak q; for a desk hiring from banks, it's often the path of least resistance. ## The comparison, condensed | | DuckDB | ClickHouse | TimescaleDB | QuestDB | kdb+ | |---|---|---|---|---|---| | Ops burden | None | Real | Moderate (it's Postgres) | Light | Real + license | | Historical scan speed | Excellent (single node) | Excellent (any scale) | Good with compression | Good | Excellent | | Live ingestion | Not its job | Good (batched) | Good | Excellent | Excellent | | As-of joins | Native | Native | Emulated | Native | Native, best-in-class | | Joins vs. reference data | Good | Limited by design | Best (full Postgres) | Basic | Good, in q | | Cost | Free | Free / cloud | Free / cloud (license nuance) | Free / cloud | Commercial | | Team fit | Solo researcher | Data-heavy team | Mixed workload team | Small live desk | Institutional desk | ## Defaults we'd actually pick - **Solo researcher, historical work:** Parquet + **DuckDB**. Revisit only when a second concurrent writer or a second machine appears. - **Small team, growing archive, shared queries:** **ClickHouse**, with Parquet remaining the interchange/archive format. - **Live capture of your own feeds with SQL on fresh data:** **QuestDB** for the hot path, exporting daily partitions to Parquet for research. - **Tick data must join operational/relational state constantly:** **TimescaleDB**, accepting the scan-speed ceiling. - **Institutional desk with q skills in-house:** **kdb+**, knowingly. ## Benchmark it yourself — in an afternoon Any comparison you didn't run yourself is a hypothesis. The minimal honest harness: one liquid symbol-month of real ticks (not synthetic — compression and cardinality behave differently), loaded identically into each candidate, timing five queries: (1) full-scan aggregate by symbol-day, (2) one-symbol one-day slice, (3) trades-to-quotes as-of join, (4) 1-minute OHLCV resample, (5) your actual signal query. Run cold and warm, record load time and on-disk size. The spread between engines on query 3 and query 5 is usually what decides it — and it's exactly what generic benchmarks don't measure. *Assessments dated September 2026. No sponsored placements; if that ever changes it will be disclosed inline.* --- **Also in this series:** [The full research stack →](/guides/quant-research-stack/) · [Choosing a data vendor →](/guides/market-data-vendors/) · **Size it first:** [tick data storage sizer →](/tools/tick-data-storage-sizer/) ---------------------------------------------------------------------- # Transaction Costs, Slippage, and Market Impact: From Paper Alpha to Tradeable Alpha URL: https://thequant.space/guides/transaction-costs-slippage-market-impact/ Section: Guides Date: 2026-09-07 Description: How to model transaction costs in backtests: the four cost components, the square-root impact law, honest cost ranges by asset class, and the turnover arithmetic that kills most published strategies. Most published strategies don't fail in the market; they fail in the spreadsheet cell the authors left empty. Costs are where academic alpha goes to die, and the arithmetic is brutally simple: **annual cost drag ≈ annual turnover × per-trade cost**. A strategy turning over 20× a year at 15 bps per side pays ~6%/yr before earning anything. That single line disqualifies more papers than any statistical test — which is why the cost question is check #3 in [our backtest framework](/guides/evaluate-trading-backtest/#3-what-happens-at-2-the-assumed-costs) and a core input to the [rigor score](/score-guide/). ## The four components, separated Lumping "costs" into one number hides which one is eating you. Keep them separate: **1. Fees and commissions** — the only component you know exactly: exchange fees, commissions, regulatory charges, borrow costs on shorts (routinely forgotten and often the largest line for crowded shorts), and funding rates for perpetuals. **2. Spread** — you cross half the bid-ask on every entry and exit as a taker. For liquid US large caps this is ~1–2 bps; for small caps, tens of bps; for long-tail crypto and [prediction markets](/guides/prediction-market-backtesting/#2-phantom-liquidity), it can exceed the entire claimed edge. Backtests filling at mid-price assume this component is zero — the single most common cost lie. **3. Market impact** — the cost of your own size moving the price. The workhorse empirical model is the **square-root law**: impact scales roughly with σ · √(Q/V) — order size Q relative to daily volume V, scaled by volatility. Its consequences are what matter: impact grows *sublinearly but relentlessly* with size; doubling capital does not halve returns, but it always lowers them; and every strategy has a **capacity** at which impact equals edge. Impact modeling done properly is a [microstructure](/topics/market-microstructure/) and [optimal-execution](/topics/high-frequency-trading/) discipline — at research stage you mostly need to know whether you're in the regime where it binds (roughly: orders above ~1–5% of ADV). **4. Timing/implementation slippage** — the drift between decision price and fill price: signal-to-order latency, partial fills, and the gap risk on stops. Detected by the same [one-bar-lag test](/guides/look-ahead-bias-point-in-time-data/#4-same-bar-execution) that exposes look-ahead — the two failures are cousins. ## Honest research-stage cost ranges Order-of-magnitude per-side costs for a *small* (impact-free) taker, September 2026; use as priors, not gospel, and verify against your own fills: | Instrument class | All-in per side | Notes | |---|---|---| | US large-cap equities | 1–5 bps | spread-dominated; near-zero commissions | | US small caps | 10–50+ bps | spread + borrow on shorts | | Major FX | 0.5–2 bps | venue-dependent | | Liquid futures | 0.5–2 bps | half-tick spreads | | BTC/ETH majors | 2–10 bps | taker fees dominate; maker rebates change the game | | Long-tail crypto | 25–200 bps | spread + thin books | | Options | wide | spread as % of premium is the killer; model per-contract | Two uses for this table. First, **fill the empty cell**: when a paper reports gross returns, apply these and see what survives — most [stat-arb](/topics/statistical-arbitrage/) and intraday results shrink dramatically. Second, **bound your own optimism**: if your backtest's assumed cost sits below the table's floor for your asset class, you're not conservative, you're wrong. ## The sensitivity curve beats the point estimate Because every number above is uncertain, the robust practice isn't a better point estimate — it's the **cost-sensitivity curve**: P&L at 1×, 2×, 3× assumed costs, required by [Gate 1 of the production checklist](/guides/production-checklist/#gate-1--validation-before-writing-any-production-code). The shape tells you what you own: a strategy losing 20% of its return at 2× costs is robust; one going negative is a cost model with delusions. Pair it with **turnover reporting** — the strategy that halves its Sharpe when you halve its turnover was mostly trading for the thrill. Two design implications fall straight out: prefer signal constructions that trade less (rebalance bands beat calendar rebalancing; signal averaging beats signal chasing), and remember that costs are the one strategy input you can *engineer* — execution quality, maker vs taker, and patience are alpha sources with none of alpha's decay. ## Live costs are the ongoing measurement After go-live, realized cost per trade — actual fills vs decision price, from your [reconciliation loop](/guides/notebook-to-production/#5-logs-metrics-and-the-alert-that-actually-matters) — becomes a first-class metric. It closes the loop: your backtest's cost assumption stops being a prior and becomes a measured, auditable number, in exactly the spirit of [assumed vs measured](/guides/auditing-a-monte-carlo/) that runs through everything we publish. --- **Where impact modeling lives:** [Market microstructure hub →](/topics/market-microstructure/) · [HFT & execution hub →](/topics/high-frequency-trading/) · **The arithmetic's next step:** [statistical vs economic significance →](/guides/statistical-vs-economic-significance/) ---------------------------------------------------------------------- # Volatility Targeting: Benefits, Hidden Leverage, and Drawdown Risk URL: https://thequant.space/guides/volatility-targeting/ Section: Guides Date: 2026-09-07 Description: Volatility targeting mechanics: why scaling by inverse vol has worked, the leverage it quietly embeds, gap risk and vol-spike deleveraging, and the estimator choices that change everything. **Not an implementation claim.** Vol targeting is a *risk-shaping layer*, not an edge; this guide covers what the layer actually does — including the leverage and procyclicality its cleaner marketing omits. **Definition.** Volatility targeting sizes exposure inversely to estimated volatility: position = target_vol / estimated_vol, so the portfolio runs at roughly constant risk. When vol is 10% and target is 15%, you're levered 1.5×; when vol spikes to 30%, you cut to 0.5×. Both halves matter: **the strategy is a leverage rule**, and its historical benefits and its failure modes are both consequences of that. ## Why it has worked: the two channels The evidence for vol-scaled versions of equity and [trend strategies](/guides/cross-sectional-vs-time-series-momentum/#mechanics-that-decide-live-behavior) improving Sharpe and drawdowns rests on two real regularities: **vol clustering** ([the most forecastable object in finance](/topics/volatility/) — high vol today predicts high vol tomorrow, so cutting after spikes avoids more of the subsequent variance than return), and the **leverage-effect correlation** — in equities, high vol associates with poor returns, so inverse-vol sizing tilts exposure toward the better-returning calm states. Honest reading of the record: the improvement is genuine, asset-class dependent (strongest in equities and equity-correlated premia, weaker where the vol-return correlation is absent), and *partly a payment for the fine print below* — insurance premia collected until the gap. ## The fine print, in mechanism order - **Hidden leverage in calm regimes**: constant-vol means maximum leverage exactly when spreads and vol are compressed — the classic pre-storm configuration. Your "12% vol" strategy is a 2× levered position wearing calm-weather clothes, with [financing and margin terms](/guides/risk-parity-min-variance/#common-ways-the-family-fails-in-production) that reprice in stress. - **Gap risk is unscaled**: sizing controls exposure to *forecastable* vol; an overnight gap hits the leveraged position at its pre-gap size. Vol targeting converts smooth-drawdown risk into [jump-drawdown risk](/guides/high-sharpe-ratio-not-investable/#3-volatility-isnt-risk--skew-is) — better average drawdowns, unchanged (or worse) tail-of-the-tail. - **Procyclical deleveraging**: cutting after spikes means selling into falling, illiquid markets alongside every other vol-targeted mandate — a crowding channel with systemic flavor: the aggregate vol-target complex amplifies exactly the moves it individually manages. - **The turnover bill**: tracking vol means [trading the whole book on vol changes](/guides/turnover-factor-returns/#common-ways-turnover-destroys-live-factor-returns); tight tracking of a fast estimator can double a slow strategy's turnover. ## The plumbing that decides live behavior Vol targeting is operationally an *estimator choice* wearing a strategy name: the estimator ([realized at what sampling, EWMA at what decay, GARCH-family](/guides/microstructure-noise-realized-volatility/)) sets responsiveness; responsiveness trades whipsaw against gap exposure; caps/floors on leverage and rebalance bands set the [turnover](/guides/turnover-factor-returns/)-tracking compromise. Two identical "12% vol target" products with different estimator half-lives are different strategies — and a paper's vol-targeting improvement that [survives only one estimator configuration](/guides/robustness-trading-strategy/#1-parameter-robustness--the-plateau-test) is a parameter spike, not a property. ## Common ways vol targeting fails in production - **The gap through the leverage**: February 2018-style vol events — levered position, discontinuous repricing, deleveraging *after* the loss. The layer worked as designed; the design's tail was misunderstood. - **Estimator regime mismatch**: an EWMA tuned on one era [over- or under-reacts in the next](/guides/regime-dependence/); the layer's parameters need the same regime scrutiny as any signal's. - **Double-scaling**: a vol-targeted strategy inside a vol-targeted portfolio compounds leverage rules nobody sized jointly. - **Backtest flattering via the vol-return correlation**: in samples where the correlation was strong, the layer looks like alpha; attribute the improvement to its channels before [crediting the wrapper](/guides/evaluate-trading-backtest/) — and check the asset classes where the channel is absent. ## Questions to ask a paper 1. How much of the improvement is the vol layer vs the underlying signal — is the unscaled version shown? 2. What estimator, what half-life, and does the result [survive the neighborhood](/guides/robustness-trading-strategy/)? 3. What happened at the sample's gap events — and are overnight moves handled as gaps or smooth paths? 4. What's the layer's own turnover and cost bill? 5. What leverage does the strategy reach in the calmest decile of the sample? ## What our scoring captures Vol-management papers rank well on [rigor](/score-guide/) when they attribute improvements to channels, show unscaled baselines, and report leverage distributions — the [volatility hub's](/topics/volatility/) top tier does all three. The score won't flag double-scaling or your portfolio's aggregate leverage — composition risks live outside any single paper — which is what [the production checklist's](/guides/production-checklist/) exposure limits exist to catch. --- **The estimator layer:** [realized vol & noise →](/guides/microstructure-noise-realized-volatility/) · **The family:** [risk parity →](/guides/risk-parity-min-variance/) · **The tail view:** [why high Sharpe ≠ investable →](/guides/high-sharpe-ratio-not-investable/) ---------------------------------------------------------------------- # Walk-Forward Analysis and Out-of-Sample Testing, Done Honestly URL: https://thequant.space/guides/walk-forward-out-of-sample-testing/ Section: Guides Date: 2026-09-07 Description: How to run out-of-sample tests that mean something: why k-fold fails on market data, purging and embargoes, anchored vs rolling walk-forward, and the evaluate-once discipline. Every quant paper claims out-of-sample results; few earn the phrase. "Out-of-sample" is not a data split — it's a **protocol**, and the protocol's entire value comes from what you *didn't* do: didn't look, didn't iterate, didn't quietly try again. This guide covers the mechanics (splits, purging, walk-forward structure) and the discipline, which is harder than the mechanics. ## Why the ML playbook fails on market data Standard k-fold cross-validation assumes exchangeable observations. Financial time series violate this three ways, each one a leak: 1. **Serial dependence**: adjacent observations share information — training on Tuesday and testing on Wednesday is nearly testing in-sample, especially with overlapping feature windows or multi-day labels. 2. **Boundary leakage**: a label looking k periods ahead, computed near a split boundary, *contains* test-period outcomes. Train/test contamination via labels is the [look-ahead family's](/guides/look-ahead-bias-point-in-time-data/#6-normalization-and-statistics-computed-on-the-full-sample) sneakiest member — we flagged exactly this (100-event labels, 9-tick window overlap at an 80/20 boundary) in [a published replication review](/flowcharts/temporal-kolmogorov-arnold-networks-t-kan-for-high-frequen/#discussion). 3. **Random shuffling destroys regimes**: shuffled folds train on 2022 to predict 2019, giving every fold knowledge of all regimes — an option live trading never has. The fixes, standard since López de Prado formalized them: **purging** (drop training samples whose labels overlap the test window) and an **embargo** (exclude a further buffer — at least feature-window + label-horizon — after each test block). If you use CV on market data at all, it's purged, embargoed, and time-ordered, or it's leaking. ## The walk-forward structure Walk-forward is the honest default because it mimics deployment: fit on the past, trade the next slice, roll forward, never look back. ```text anchored: [train ........][test] [train ..............][test] rolling: [train ....][test] [train ....][test] ``` - **Anchored** (expanding window) uses all history — more data, slower adaptation, implicitly assumes old regimes stay relevant. - **Rolling** (fixed window) adapts and forgets — the window length becomes a *hyperparameter with regime-size meaning*, and results that hinge on it deserve the same [plateau-vs-spike scrutiny](/guides/evaluate-trading-backtest/#5-does-it-survive-the-neighborhood-of-its-own-parameters) as any other parameter. - Either way, **re-optimize inside each training window only** — parameters chosen once on the full sample then "walk-forward tested" is in-sample fitting wearing a lab coat. What to read from the output: not just aggregate performance but **stability across windows** (one golden window carrying the average is a regime bet, not an edge — check the subperiod table), and **parameter drift** across refits (parameters that lurch every window are chasing noise). The [ML](/topics/machine-learning/) and [RL](/topics/reinforcement-learning/) literatures fail these two checks more than any others. ## The evaluate-once rule Here is the uncomfortable truth that makes protocols matter: **a holdout is spent the moment it influences a decision.** Evaluate on the holdout, tweak, evaluate again — after a few cycles the holdout is training data with ceremony, and its Sharpe is quietly the [maximum of several tries](/guides/deflated-sharpe-ratio/). The discipline: - Decide the final test *protocol and success criteria in writing* before touching the holdout ([Gate 1](/guides/production-checklist/#gate-1--validation-before-writing-any-production-code) requires exactly this). - Spend it once. If the result fails and you keep iterating, **carve a new, never-touched period** for the next final exam — and count the spent one in your [trial count](/guides/statistical-vs-economic-significance/). - Log every evaluation. An honest lab notebook of holdout touches is the cheapest integrity infrastructure that exists. ## The out-of-sample tests nobody can fake Two sources of OOS data are immune to your own iteration. **Post-publication data** for published strategies: a 2015 paper meets 2016–2026 markets it has never seen — the cleanest test in the field and the heart of [reproduction work](/guides/reproduce-quant-research/). And **the future**: paper trading through your real execution path, then minimum-size live trading, per the [go-live protocol](/guides/production-checklist/#gate-6--go-live-protocol). Every backtest, however disciplined, shares your assumptions about fills and costs; the [fill-quality gap](/guides/transaction-costs-slippage-market-impact/#live-costs-are-the-ongoing-measurement) between simulation and live is itself the final out-of-sample statistic. ## The checklist - ☐ Time-ordered splits only; no shuffling - ☐ Purge overlapping labels; embargo ≥ window + horizon at every boundary - ☐ All transforms fit inside training windows (scalers included) - ☐ Re-optimization inside each walk-forward window - ☐ Per-window results and parameter paths reported, not just the aggregate - ☐ Holdout criteria written down first; holdout spent once; touches logged - ☐ Trial count carried into [significance judgment](/guides/statistical-vs-economic-significance/) --- **The multiple-testing arithmetic:** [Deflated Sharpe ratio →](/guides/deflated-sharpe-ratio/) · **Where leakage hides:** [look-ahead taxonomy →](/guides/look-ahead-bias-point-in-time-data/) ---------------------------------------------------------------------- # What 'Robustness' Actually Means in a Trading Strategy URL: https://thequant.space/guides/robustness-trading-strategy/ Section: Guides Date: 2026-09-07 Description: A precise definition of strategy robustness: the five dimensions (parameters, universe, time, costs, implementation), how to test each, and how to read a paper's robustness section adversarially. **Definition.** A strategy is robust to the degree that its performance **degrades gracefully under perturbations it was not optimized against** — in parameters, universe, time, costs, and implementation details. Two words carry the weight: *gracefully* (robustness is a slope, not a pass/fail), and *not optimized against* (surviving checks you tuned on is [selection](/guides/backtest-overfitting/), not robustness). The word appears in nearly every empirical paper; the property appears in far fewer. ## The five dimensions, each with its test **1. Parameter robustness — the plateau test.** Perturb every tunable ±20–50% and map the surface. The standard: a *plateau* where performance declines smoothly, not a [spike](/guides/evaluate-trading-backtest/#5-does-it-survive-the-neighborhood-of-its-own-parameters). Numerical rule of thumb: if the best cell's Sharpe is more than ~1.5× the median of its neighbors, you've found noise's address, not an edge. **2. Universe robustness — the transfer test.** Same rule, different assets: other size buckets, sectors, countries, or venues. Perfect transfer isn't expected — mechanisms have habitats — but the *mechanism story must predict the transfer pattern*. Value that works everywhere except where the paper's story says it shouldn't: robust. Momentum that works only in the 200 stocks the authors kept: [sculpted](/guides/p-hacking-financial-research/#the-seven-practices-in-ascending-order-of-self-deception). **3. Time robustness — subperiods and regimes.** Split by decades, volatility regimes, and rate environments. The standard again is graceful: a strategy earning in 70% of years with understood losses beats one carried entirely by [a single golden window](/guides/walk-forward-out-of-sample-testing/#the-walk-forward-structure). Where an edge is *legitimately* regime-dependent, robustness means the paper says so and identifies the regime — the subject of [its own guide](/guides/regime-dependence/). **4. Cost robustness — the sensitivity curve.** P&L at 1×/2×/3× assumed costs ([the standing requirement](/guides/transaction-costs-slippage-market-impact/#the-sensitivity-curve-beats-the-point-estimate)). Graceful: 2× costs takes a proportionate bite. Fragile: 2× flips the sign. **5. Implementation robustness — the sloppiness test.** Delay signals a bar, execute at open instead of close, round positions, miss 5% of trades at random. Real mechanisms survive sloppy hands with degraded but recognizable performance; artifacts require *exact* execution, because [the exactness is where the leak lives](/guides/look-ahead-bias-point-in-time-data/#4-same-bar-execution). ## The example that compresses the concept Two moving-average strategies, same backtest Sharpe of 1.4. Strategy A: lookback grid 40–70 all score 1.1–1.5; works on 8 of 10 futures; loses moderately in 2018; survives 2× costs at SR 0.9. Strategy B: lookback 52 scores 1.4, lookbacks 45 and 60 score 0.2; works on 2 markets; all profit from 2020; dies at 1.5× costs. Identical headline, opposite objects: A is an edge with error bars, B is a coordinate in noise. Every robustness dimension is a projection of that same distinction. ## Reading a paper's robustness section adversarially - **Count the casualties.** Honest grids have edges; a section where [every check passes](/guides/p-hacking-financial-research/#tells-a-reader-can-spot-in-the-published-paper) was curated. The most credible sentence in empirical finance is "the effect disappears when…". - **Check what's *missing*:** the dimension not tested is usually the one that fails. No cost sensitivity in a high-turnover paper is a confession by omission. - **Distinguish robustness from re-optimization:** "results hold with parameters re-tuned per subperiod" tests the *pipeline's* ability to fit, not the strategy's stability. - **Beware robustness-by-average:** a table averaging over specifications can hide that half were negative. ## The checklist - ☐ Parameter surface mapped; plateau confirmed - ☐ Transfer tested; pattern matches the mechanism story - ☐ Subperiod/regime table with losses explained, not excused - ☐ Cost curve at 1×/2×/3× - ☐ One-bar-delay and sloppy-execution variants - ☐ At least one reported check that *binds* ## What our scoring captures — and misses The presence and breadth of real robustness sections is one of the strongest inputs to our [empirical-rigor score](/score-guide/) — it separates Street Traders from Holy Grail papers more reliably than any single metric. The structural miss: we score the checks a paper *ran*; the five-dimension framework tells you which checks it *owed*. The gap between those two lists is your remaining risk, and [reproduction](/guides/reproduce-quant-research/#stage-5--stress-the-part-the-paper-didnt-do) is how you close it. --- **Companions:** [Backtest overfitting →](/guides/backtest-overfitting/) · [Regime dependence →](/guides/regime-dependence/) · [The eight backtest questions →](/guides/evaluate-trading-backtest/) ---------------------------------------------------------------------- # What Is Backtest Overfitting? Definition, Detection, and Defenses URL: https://thequant.space/guides/backtest-overfitting/ Section: Guides Date: 2026-09-07 Description: Backtest overfitting explained: how selection among many trials fits noise, a concrete numerical demonstration, the PBO/deflated-Sharpe detection tools, and the defenses that actually work. **Definition.** Backtest overfitting is when a strategy's historical performance reflects **fitting the sample's noise rather than a persistent mechanism** — almost always through *selection*: many variants evaluated, the best kept, the trial count forgotten. It requires no bad faith and no complex model; a grid of moving-average lengths overfits as efficiently as a neural net. It is the central quality problem in strategy research, and the reason "it worked in the backtest" is the beginning of an argument, not the end. ## The ten-line demonstration Replicable in any language, or visually in our [equity-curve simulator](/tools/equity-curve-simulator/): ```python import numpy as np rng = np.random.default_rng(0) rets = rng.normal(0, 0.01, (252, 1000)) # 1,000 strategies, zero true edge sharpes = rets.mean(0) / rets.std(0) * 252**0.5 print(sharpes.max()) # ≈ 3.2 — from pure noise ``` One year of daily data, a thousand coin-flip strategies: the best prints an annualized Sharpe above 3. Nothing was learned about markets; something was learned about maxima — the expected best of N grows like the [extreme-value formula](/guides/deflated-sharpe-ratio/#step-2-the-maximum-of-n-tries-is-not-a-fair-draw) says it must (≈2.5 at N=100, ≈3.25 at N=1,000). Every parameter sweep you run is this experiment with better branding. ## How it happens in practice, ranked by frequency 1. **Parameter selection** — the grid search that ends at the best cell. 2. **Feature/universe/period selection** — trying signals, asset lists, and start dates until one works ([sample choice is a trial too](/guides/how-to-read-quant-finance-papers/)). 3. **Holdout erosion** — the "out-of-sample" set [evaluated repeatedly](/guides/walk-forward-out-of-sample-testing/#the-evaluate-once-rule) until it agrees. 4. **Meta-overfitting** — trying validation *methods* until one validates. Yes, this happens. Note what's *not* on the list: model complexity. A 2-parameter strategy selected from 500 tries is more overfit than a 500-parameter model evaluated once. **Overfitting lives in the selection process, not the parameter count.** ## Detection: the standing toolkit - **[Deflated Sharpe ratio](/guides/deflated-sharpe-ratio/)** — is the best result better than the best-of-N-by-luck benchmark? - **[PBO via CSCV](/guides/cscv-explained/)** — across many symmetric in/out splits, how often does the in-sample winner underperform out-of-sample? A direct probability-of-overfitting estimate. - **Parameter-plateau inspection** — [spike vs plateau](/guides/evaluate-trading-backtest/#5-does-it-survive-the-neighborhood-of-its-own-parameters); noise fits are spiky. - **Performance-decay slope** — overfit strategies degrade from in-sample → holdout → [paper trading](/guides/walk-forward-out-of-sample-testing/#the-out-of-sample-tests-nobody-can-fake) monotonically; the slope is the tell. ## Defenses, ranked by what they cost you | Defense | Cost | Effect | |---|---|---| | Log every trial ([experiment log](/guides/quant-research-pipeline/#1-idea-intake--the-experiment-log)) | An afternoon | Makes N honest — enables everything below | | Pre-registered success criteria | Discipline | Kills post-hoc goalpost moves | | Economic-mechanism requirement | Ideas rejected | Strategies must have a *reason*; noise rarely does | | Purged walk-forward + evaluate-once holdout | Data | The [protocol](/guides/walk-forward-out-of-sample-testing/) | | DSR/PBO gates before capital | Some winners rejected | The statistical backstop | ## Questions to ask when reading a paper 1. What is the visible search space (tables, appendix grids) — and what multiple of it is the plausible *invisible* one? 2. Is there a mechanism story that predated the result, or a result in search of a story? 3. Does the paper report anything like PBO, DSR, or even trial counts? (Rare — and instantly credibility-raising.) 4. How does performance change in the years *after* the sample ends? ## What our scoring captures — and misses Overfitting is the [rigor axis's](/score-guide/) hardest target: protocol quality (OOS structure, robustness sections) is visible in a paper's text and scored, but selection intensity is invisible by nature — the 2,000 discarded variants leave no textual trace. Treat our score as evidence the *reported* experiment was run well, and apply the trial-count correction yourself. The papers that make this easy — disclosed search spaces, DSR-style adjustments — are the ones worth [replicating first](/guides/choose-papers-to-replicate/). --- **The math:** [Deflated Sharpe →](/guides/deflated-sharpe-ratio/) · [CSCV/PBO →](/guides/cscv-explained/) · **The process fix:** [research pipeline →](/guides/quant-research-pipeline/) · **Run your numbers:** [deflated Sharpe calculator →](/tools/deflated-sharpe-calculator/) ---------------------------------------------------------------------- # What Makes a Factor Tradable? URL: https://thequant.space/guides/what-makes-a-factor-tradable/ Section: Guides Date: 2026-09-07 Description: The gap between a published factor premium and a tradable one: implementation costs, capacity, crowding, borrow reality, and the checklist that converts paper premia into honest expectations. **Not an implementation claim.** This guide is the translation layer between [factor papers](/guides/evaluate-factor-investing-paper/) and portfolios — the arithmetic that determines *how much* of a published premium survives, not an endorsement that any particular factor's residual is worth harvesting. **Definition.** A factor is *tradable* to the extent its paper premium survives five sequential haircuts. Each gate below has a magnitude attached; multiply them through before any published number enters your expectations. ## Gate 1: Implementation costs — the turnover haircut The arithmetic from [the turnover guide](/guides/turnover-factor-returns/): annual drag = turnover × per-side cost × 2. Value-family factors (annual turnover ~20–50%) lose tens of bps; momentum-family (100–300%+) can lose [several percent](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges) — against premia measured in low single digits. The construction detail that decides it: rebalance frequency and bands are *choices*, and the tradable version of most factors is a deliberately slowed version of the published one. ## Gate 2: Capacity — where does the premium live? Decompose the paper's sort by size bucket. A premium earned mostly in the bottom size quintile has the [capacity of that quintile](/guides/capacity-constraints/) — often single-digit millions per name of realistic daily trading. The [zoo literature's finding](/guides/interpret-factor-zoo-paper/#what-survives-everyones-methodology) that micro-caps carry a large share of anomalies is, restated, a finding that much of the zoo is capacity-free. Value-weighted, large-cap-only replication is the honest capacity test — [and the harsh one](/guides/evaluate-factor-investing-paper/#the-factor-paper-checklist). ## Gate 3: The short side — half the premium, most of the problems Published long-short premia often earn disproportionately from the short leg — which is where implementation reality concentrates: borrow fees (routinely 1–10%+ annualized on the crowded shorts every factor selects), recalls and buy-ins at the worst moments, and short-squeeze dynamics that turn the leg's [negative skew](/guides/high-sharpe-ratio-not-investable/#3-volatility-isnt-risk--skew-is) structural. The standard translations: long-only tilts capture some premia at a fraction of the paper spread; 130/30-style constructions split the difference. A factor whose premium *requires* the full short leg is tradable only at institutional shorting infrastructure. ## Gate 4: Crowding — the premium's reflexive component A published factor is a *shared* factor: its holders become a correlated bloc whose entries compress the premium and whose exits create the factor's own crash risk (the quant unwinds of 2007 and the momentum reversals of 2009/2020 are the exhibits). Crowding is hard to measure and impossible to ignore; proxies — valuation spreads of the factor portfolio, short interest on the short leg, factor-flow estimates — belong on the same monitor as the premium itself. Crowding also links gates: it *raises* effective costs (everyone rebalances together) and *shrinks* capacity (the trade's aggregate size is the constraint, not yours). ## Gate 5: Decay — the premium's time derivative [Post-publication decay](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer) of a third to a half is the base rate. The tradable expectation is the *recent, net* premium — which for several famous factors is statistically indistinguishable from zero, per the [replication meta-studies](/guides/interpret-factor-zoo-paper/). Factors with structural or risk-based mechanism stories decay slower than pure statistical regularities — one of the few places where [the mechanism question](/guides/backtest-overfitting/#defenses-ranked-by-what-they-cost-you) pays cash. ## The worked example Paper: 6%/yr long-short premium, equal-weighted, monthly rebalanced, 1980–2015. Translation: value-weight and drop micro-caps (−2%), slow to quarterly with bands (turnover 150%→60%, saving 1.1% of a 1.35% cost drag), long-only tilt (capture ~40% of remaining spread), decay haircut (−35%): **≈1.0–1.3%/yr of tradable tilt premium** — before crowding risk. Not nothing at scale; not the paper's 6%; and exactly the arithmetic no abstract performs. ## Common ways factor implementations fail in production - **Tracking-error boredom**: the tilt underperforms its benchmark for [statistically normal multi-year stretches](/guides/confidence-intervals-strategy-research/#the-numbers-nobody-internalizes); capital's patience, not the premium, is what breaks. - **Rebalance-date crowding**: trading the same month-end as every indexed factor product. - **Definition drift**: the live signal (vendor fundamentals, [restatement handling](/guides/look-ahead-bias-point-in-time-data/)) quietly differs from the paper's CRSP/Compustat construction. - **The factor works; your version doesn't**: neutralizations and constraints [reshape the bet](/guides/statistical-arbitrage-signal-to-portfolio/#step-1-neutralization--deciding-what-youre-betting-on) until realized returns decorrelate from the published series. ## Questions to ask 1. What's the premium under value-weighting, ex-micro-caps, net of the paper's own turnover? 2. Long leg, short leg, or spread — where does it actually accrue? 3. What's the borrow bill on the short decile *today*? 4. What has the factor returned since publication? ## What our scoring captures The gates are the rigor axis's practical half: papers reporting net-of-cost, capacity-bucketed, post-sample results occupy the top of the [factor hub](/topics/factor-investing/), and the [quadrant system](/score-guide/) separates premium-documentation (academic) from premium-harvesting (implementable) papers cleanly. What no score provides is gates 4–5's *current* state — crowding and decay are live variables; the archive dates every paper so you can apply the clock yourself. --- **The haircut arithmetic:** [turnover →](/guides/turnover-factor-returns/) · [capacity →](/guides/capacity-constraints/) · **The meta-evidence:** [factor-zoo papers →](/guides/interpret-factor-zoo-paper/) ---------------------------------------------------------------------- # When to Use Parquet, Postgres, or a Columnar Database URL: https://thequant.space/guides/parquet-vs-database/ Section: Guides Date: 2026-09-07 Description: The three storage archetypes for quant research — files, relational, columnar-analytical — matched to workload shapes, with the two-tier default and the migration triggers. The [engine comparison](/guides/tick-data-databases/) answers "which analytical database?"; this guide answers the question *before* it: **does this data belong in files, a relational database, or a columnar engine at all?** The three are archetypes with different contracts, and most storage regret comes from putting a workload in the wrong archetype, not the wrong product. ## The three contracts **Files (Parquet + DuckDB/Polars)** promise: maximal scan throughput per dollar, zero operations, universal tool compatibility, [cheap archival](/guides/store-tick-data-efficiently/). They decline: concurrent writers, transactions, fine-grained updates, enforcement of anything. Files are *immutable facts in bulk* — which is precisely what [research reads](/guides/quant-research-stack/#layer-2-storage--parquet--duckdb-until-it-hurts) look like. **Relational (Postgres)** promises: transactions, constraints, concurrent readers *and* writers, point lookups and joins, [a schema that enforces itself](/guides/notebook-to-production/#2-postgres-is-the-source-of-truth-not-your-process-memory). It declines: full-table analytical scans at columnar speeds. Postgres is *mutable truth with rules* — orders, positions, [experiment logs](/guides/quant-research-pipeline/#1-idea-intake--the-experiment-log), the [bitemporal reference layer](/guides/point-in-time-data-architecture/), anything where `UNIQUE` doing its job [saves an account](/guides/notebook-to-production/#2-postgres-is-the-source-of-truth-not-your-process-memory). **Columnar-analytical (ClickHouse-class)** promises: file-like scan speeds with a serving layer — concurrent analytical queries over tens of billions of rows, continuous ingestion, [materialized aggregation](/guides/tick-data-databases/#clickhouse--the-scale-up-column-store). It declines: transactional semantics, cheap mutation, operational simplicity. It's *files that answer queries while being written* — a contract you need later than you think. ## The workload → archetype map | Workload | Archetype | Why | |---|---|---| | Backtest scans over history | Files | Read-only bulk; [scan speed is the product](/guides/store-tick-data-efficiently/) | | Orders, fills, positions, run-state | Postgres | Constraints and transactions are the point | | [PIT reference & fundamentals](/guides/point-in-time-data-architecture/) | Postgres | Bitemporal queries, joins, moderate size | | Feature store (computed, versioned) | Files | Regenerable bulk; [version the code](/guides/versioning-datasets-backtests/) | | Live tick capture + intraday queries | Columnar (or [QuestDB-style](/guides/tick-data-databases/#questdb--ingestion-first-time-series)) | Ingest-while-querying is the defining need | | Team-shared analytical layer | Columnar | Concurrent scans files handle badly | | [Subscriber/event tracking, app state](/guides/notebook-to-production/) | Postgres/D1-class | Small, transactional, boring | ## The two-tier default For a [one-to-three-person shop](/guides/quant-research-stack/): **Parquet files for everything bulky and immutable; one Postgres for everything mutable and precious.** No third tier until a trigger fires. This split also has a clean [backup story](/guides/notebook-to-production/#6-backups-you-have-restored-at-least-once) — `pg_dump` for the precious tier, object-storage sync for the bulk tier — and a clean [reproducibility story](/guides/docker-reproducible-research/): files are content-addressable; databases are dumps. ## Migration triggers — specific, not vibes To columnar, when: a second person's queries contend with yours daily; you're ingesting continuously *and* querying the same data intraday; single-node scans exceed coffee-length despite [layout fixes](/guides/store-tick-data-efficiently/#the-four-decisions-that-matter); or the working set outgrows one machine's NVMe. To bigger-Postgres/managed, when: [ops time exceeds its cost](/guides/notebook-to-production/), not before. The anti-trigger worth naming: **"we might need it later" is not a trigger** — the [stack guide's rule](/guides/quant-research-stack/#what-it-costs-september-2026-order-of-magnitude) that infrastructure is bought in response to bottlenecks applies doubly to distributed databases, whose operational bill arrives instantly while their benefits wait for scale you may never reach. ## Common archetype mistakes - **Ticks in Postgres**: row storage, WAL amplification, and vacuum meet billions of rows; scans crawl, the instance swells — the most common first-timer error. - **Positions in Parquet**: mutable state in immutable files → last-writer-wins corruption and no constraints; the [idempotency machinery](/guides/production-checklist/#gate-4--execution-safety-the-account-savers) has nothing to grip. - **A cluster for a laptop-sized problem**: the columnar tier's admin tax paid on data DuckDB scans in seconds. - **The hybrid nobody versioned**: features half in files, half in tables, no [lineage](/guides/versioning-datasets-backtests/) — reproducibility dies in the seam. ## Questions to ask before placing any dataset 1. Is it mutable truth or immutable fact? 2. Who writes, who reads, how concurrently? 3. What's the query shape — scan, lookup, or as-of join? 4. Which tier's [backup/restore story](/guides/notebook-to-production/#6-backups-you-have-restored-at-least-once) does it need? ## Where this connects This is the decision *above* the [engine comparison](/guides/tick-data-databases/) — read that one when a trigger fires and the columnar tier is genuinely due. The two-tier default is the storage half of the [research stack](/guides/quant-research-stack/) and the [production stack](/guides/notebook-to-production/); the [PIT architecture](/guides/point-in-time-data-architecture/) lives in the Postgres tier by design. --- **The engines:** [database comparison →](/guides/tick-data-databases/) · **The file tier done right:** [tick storage →](/guides/store-tick-data-efficiently/) · **The versioning seam:** [datasets & backtests →](/guides/versioning-datasets-backtests/) ---------------------------------------------------------------------- # Why a High Sharpe Ratio May Not Be Investable URL: https://thequant.space/guides/high-sharpe-ratio-not-investable/ Section: Guides Date: 2026-09-07 Description: Seven reasons a high backtested Sharpe ratio fails to translate into investable returns: selection, capacity, skew, leverage limits, correlation timing, costs, and career horizons. **Definition.** The Sharpe ratio — mean excess return over volatility — measures *statistical efficiency*, not investability. "High Sharpe" and "makes money for you" are connected by a chain with seven links, and any one can fail while the Sharpe stays technically true. This is the checklist for the gap, and it applies double to every paper in the archive claiming SR > 2. ## The seven failure links **1. It's a maximum, not a draw.** Before anything else: a high Sharpe selected from many attempts is [expected under the null](/guides/deflated-sharpe-ratio/) — best-of-100 on a year of data prints ≈2.5 from noise. Everything below assumes the Sharpe survived that filter; most never get that far. **2. It doesn't scale.** Sharpe is size-invariant; returns aren't. A 3-Sharpe strategy earning 4%/yr on $500k capacity is a nice salary, not a fund — and [impact's square-root law](/guides/transaction-costs-slippage-market-impact/#the-four-components-separated) guarantees the Sharpe itself degrades as size grows. The investable statistic is **dollar alpha at achievable size**, which a ratio cannot express. High-frequency strategies are the extreme: [SR 5+ with tiny capacity](/topics/high-frequency-trading/) is common and uninvestable for outside capital. **3. Volatility isn't risk — skew is.** Sharpe's denominator treats a short-vol premium harvest and a balanced strategy identically until the cliff. Strategies that [manufacture smooth-then-catastrophic profiles](/guides/deflated-sharpe-ratio/#step-1-a-measured-sharpe-is-a-noisy-estimate) — selling tails, providing liquidity into crashes, carrying pegs — post high Sharpes *because* the risk hides in moments the ratio ignores. Check skew, kurtosis, and worst-month before admiring the ratio; the PSR formula does this arithmetically. **4. You can't always lever the efficiency into returns.** The textbook says lever the high-Sharpe/low-return strategy up. Reality bills for it: margin limits, borrow availability, [volatility-targeting's hidden gap risk](/guides/volatility-targeting/), and financing costs that scale with leverage while the edge doesn't. An unleverageable 2-Sharpe earning 3%/yr loses to a leverageable 1-Sharpe earning 12%. **5. Correlation arrives exactly when it matters.** A standalone Sharpe ignores *when* the strategy loses. Many premia — carry, liquidity provision, short-vol — earn steadily and then lose precisely in the states where your portfolio, your investors, and your margin desk are all bleeding. Conditional correlation in stress is the investability statistic; unconditional correlation flatters it. **6. Costs and slippage are convex in urgency.** The backtest's [cost assumption](/guides/transaction-costs-slippage-market-impact/) was calibrated to calm markets and patient execution. The strategy's worst days are fast markets — where spreads triple and the [fill-quality gap](/guides/notebook-to-production/#5-logs-metrics-and-the-alert-that-actually-matters) yawns open. Sharpe measured on close-to-close fills quietly assumes liquidity that disappears on the days that define the drawdown. **7. Drawdown duration vs capital patience.** Even a true, robust SR 1.5 spends [years-long stretches underwater with meaningful probability](/tools/equity-curve-simulator/) — run the fan and look at the 5th percentile path. Investability includes the sociology: will you (or your investors, or your risk committee) still be funding the strategy at month 30 of a statistically-normal drawdown? A Sharpe that outlives its capital's patience was never investable, just admirable. ## The checklist - ☐ DSR-adjusted for [trials](/guides/deflated-sharpe-ratio/) - ☐ Capacity stated; dollar alpha at that capacity computed - ☐ Skew/kurtosis/worst-month reported; short-vol content identified - ☐ Leverage required for target returns actually available at quoted financing - ☐ Stress-conditional correlation to your existing book - ☐ Costs modeled *in the states where losses occur* - ☐ Drawdown-duration distribution vs realistic capital patience ## Questions to ask when a paper leads with its Sharpe 1. At what AUM does this Sharpe halve? 2. What does the return distribution's left tail look like, and what economic state produces it? 3. Is the ratio computed on gross, net, or [close-to-close fantasy fills](/guides/evaluate-trading-backtest/)? 4. Sharpe over what period — and [what happened after](/guides/statistical-vs-economic-significance/#economic-significance-the-four-questions-t-stats-cant-answer)? ## What our scoring captures — and misses Papers reporting capacity, skew, and net-of-cost results score higher on [rigor](/score-guide/) — the axis was built to reward exactly these disclosures. What no paper score can know is link 7: investability relative to *your* capital structure and patience is a property of the reader, not the research. The ratio ranks papers; the seven links decide allocations. --- **Companions:** [Capacity constraints →](/guides/capacity-constraints/) · [Volatility targeting's fine print →](/guides/volatility-targeting/) · **See the drawdown fan:** [equity-curve simulator →](/tools/equity-curve-simulator/) ---------------------------------------------------------------------- # Why Turnover Can Destroy Factor Returns URL: https://thequant.space/guides/turnover-factor-returns/ Section: Guides Date: 2026-09-07 Description: The turnover arithmetic that decides factor profitability: signal decay vs trading cost, the rebalance-frequency trade-off, turnover-reduction techniques, and the reporting gap in academic papers. **Not an implementation claim.** This is the cost side of factor investing's ledger — the arithmetic determining how much of any [documented premium](/guides/what-makes-a-factor-tradable/) survives contact with a broker. **The one-line model.** Annual cost drag = **annual turnover × per-side cost × 2** (each unit of turnover is a sell and a buy). A momentum portfolio turning over 250%/yr at 15 bps per side pays **75 bps/yr**; the same construction in small caps at 40 bps pays **2%/yr** — against premia [measured in low single digits, pre-decay](/guides/interpret-factor-zoo-paper/#what-survives-everyones-methodology). Turnover isn't a detail of factor investing; it's the second term in its profit equation. ## The real trade-off: signal decay vs trading cost Slowing rebalancing cuts costs *and* trades on stale signals — the optimization every implementable factor solves. The governing quantity is the **signal's autocorrelation/half-life**: value-type signals (half-lives of quarters-to-years) lose little to quarterly rebalancing; fast reversal signals (half-lives of days) die if slowed. The clean framing: gross premium captured is a rising-then-flat function of rebalance frequency, cost is linear in it — the optimum sits where *marginal* signal capture equals *marginal* cost, which for most equity factors lands far slower than the monthly convention academic papers inherited. This is also why identical-signal papers disagree: [rebalance frequency is a construction choice](/guides/evaluate-factor-investing-paper/#the-worked-example-construction-is-the-hidden-factor) doing unacknowledged work. ## The four turnover-reduction techniques, and their prices 1. **Rebalance bands** (trade only past a drift threshold): typically cuts turnover 30–60% for minor tracking cost; the first tool, always. 2. **Signal smoothing** (average the signal over a window before sizing): trades noise for lag — safe for slow signals, [lethal for fast ones](/guides/statistical-arbitrage-signal-to-portfolio/#step-4-turnover-control--the-edges-tax-rate). 3. **Buy/hold rank asymmetry** (enter top decile, exit only below the top quartile): hysteresis that kills churn at the portfolio's edge, where most factor turnover lives — rank-boundary crossings, not conviction changes. 4. **Turnover-penalized optimization** (a cost term in the [optimizer](/guides/portfolio-optimization-estimation-error/)): the principled version, with the usual estimation-error caveats. Each technique changes the realized factor: the slowed portfolio's returns decorrelate somewhat from the paper's series — a cost worth paying and worth *measuring*, since it's also how [your implementation stops matching the published premium](/guides/what-makes-a-factor-tradable/#common-ways-factor-implementations-fail-in-production). ## The reporting gap to check in every paper The genre's chronic omission pattern, in ascending order of sin: turnover unreported (assume the construction's worst); turnover reported, costs unapplied ("gross of transaction costs" doing heroic work in a footnote); costs applied at a single flattering estimate (5 bps on small caps); costs applied only in the conclusion's robustness paragraph, after every headline number was gross. The [rigor standard](/guides/evaluate-trading-backtest/#3-what-happens-at-2-the-assumed-costs): turnover in the main table, net results at [honest per-class costs](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges), and a sensitivity curve. Papers meeting it are a minority — and a shortlist worth having. ## Common ways turnover destroys live factor returns - **Costs are convex at the rebalance**: everyone running the factor [trades the same names the same week](/guides/what-makes-a-factor-tradable/#gate-4-crowding--the-premiums-reflexive-component); realized costs at rebalance dates exceed calm-market estimates. - **The turnover you didn't model**: index reconstitutions, [corporate actions](/guides/corporate-actions-adjusted-prices/), delistings, and flows force trades outside the signal's schedule. - **Vol-scaling multiplies turnover**: adding a [volatility-targeting layer](/guides/volatility-targeting/) trades the *whole book* on vol changes — its turnover bill lands on top of the signal's. - **Tax friction** (taxable accounts): short-term realization can exceed trading costs entirely — the constraint that makes slow constructions dominant outside institutions. ## Questions to ask 1. What's the annual turnover, and what does it cost at [my tier's rates](/guides/transaction-costs-slippage-market-impact/#honest-research-stage-cost-ranges)? 2. What's the signal's half-life, and is the rebalance frequency justified by it or by convention? 3. Does the paper's net result survive 2× costs? 4. How much premium survives the slowest construction that still captures the signal? ## What our scoring captures Turnover disclosure is one of the [rigor axis's](/score-guide/) most mechanical checks — its absence in a performance-claiming paper caps the score, full stop. The [factor hub's](/topics/factor-investing/) upper tier is, not coincidentally, papers whose cost sections could be handed to a trader unedited. --- **The cost side:** [transaction costs →](/guides/transaction-costs-slippage-market-impact/) · **The size side:** [capacity →](/guides/capacity-constraints/) · **The full translation:** [factor tradability →](/guides/what-makes-a-factor-tradable/) ---------------------------------------------------------------------- # Deflated Sharpe Ratio & Minimum Track Record Calculator URL: https://thequant.space/tools/deflated-sharpe-calculator/ Section: Tools Description: Interactive deflated Sharpe ratio calculator: probabilistic Sharpe ratio, expected maximum Sharpe under N zero-skill trials, DSR, and minimum track record length — with skew, kurtosis, and sample length. Runs in your browser. A Sharpe ratio is an estimate with sampling error, and the best of N backtests is the *maximum* of N noisy estimates. This calculator applies the Bailey & López de Prado corrections for both: the **probabilistic Sharpe ratio** (PSR) for sample length and fat tails, and the **deflated Sharpe ratio** (DSR) for selection among trials. Every default is editable and dated **October 2026**; the reasoning is in the [deflated Sharpe ratio guide](/guides/deflated-sharpe-ratio/). Everything runs in your browser; nothing is uploaded.

Your backtest

Results

Deflated Sharpe ratio — P(true SR > luck ceiling)

—

E[max SR] under N zero-skill trials

—
annualized luck ceiling

Probabilistic Sharpe ratio vs SR*

—

Minimum track record length to beat SR*

—
—

Sensitivity to the trial count

Trials NE[max SR] (ann.)Deflated SharpeReading
Assumptions (edit me — defaults dated October 2026)

Units: the formulas run in per-period units. Annualized inputs are divided by √f (f = observations per year); T = years × f. Skewness and kurtosis are taken as measured on the per-period return series. Confidence levels for MinTRL are fixed at 95% (Z = 1.645) and 99% (Z = 2.326), one-sided. Trials are assumed independent — correlated trials have a smaller effective N, so the honest shortcut is to count distinct strategy families at full weight.

## How to read this **Probabilistic Sharpe ratio.** PSR(SR*) = Φ((SR − SR*)·√(T−1) / √(1 − γ₃·SR + ((γ₄−1)/4)·SR²)), computed in per-period units: the probability that the true Sharpe exceeds the benchmark SR*, given that your estimate came from T observations with the skew and kurtosis you entered. Negative skew and fat tails widen the estimate's error bars, so the same Sharpe earns a lower PSR. **Expected maximum Sharpe.** If you evaluate N variants with zero true skill and keep the best, you have sampled the maximum of N noisy estimates. Extreme-value theory gives E[max SR] ≈ σ_SR·[(1−γ)·Z(1−1/N) + γ·Z(1−1/(N·e))] with γ ≈ 0.5772. The default σ_SR is the standard error of a Sharpe at your sample length (about 1.0 annualized for one year of daily data), which is why 100 trials on one year of daily data produce a best-by-luck Sharpe near 2.5. The growth is logarithmic in N, so the answer is not sensitive to getting N exactly right — but it is very sensitive to pretending N was 1. **Deflated Sharpe ratio.** The PSR evaluated with SR* set to that luck ceiling. A DSR of 95%+ means your result is unlikely to be the best of a random search of the size you ran; a DSR near 50% means it is a coin flip; a DSR well below 50% means a zero-skill search would be *expected* to do as well or better. **Minimum track record length.** MinTRL = 1 + (1 − γ₃·SR + ((γ₄−1)/4)·SR²)·(Z_α/(SR − SR*))² periods: how long a track record you need before the PSR against SR* reaches the chosen confidence. Shown in years for 95% and 99%, plus the length needed to beat the luck ceiling itself. If SR ≤ SR*, no track record length suffices. What the calculator does not fix: [survivorship in the data](/guides/survivorship-bias/), [look-ahead in the pipeline](/guides/look-ahead-bias-point-in-time-data/), or unmodelled costs — a leaked backtest deflates beautifully and is still fiction. Read the [deflated Sharpe guide](/guides/deflated-sharpe-ratio/) for the worked example, the [backtest overfitting guide](/guides/backtest-overfitting/) for why selection is the central problem, and the [CSCV/PBO explainer](/guides/cscv-explained/) for the complementary probability-of-overfitting estimate. To *see* the maximum-of-N effect rather than compute it, run 100 zero-edge paths in the [equity-curve simulator](/tools/equity-curve-simulator/) and look at the best one. *Formulas: Bailey & López de Prado, "The Sharpe Ratio Efficient Frontier" (2012) and "The Deflated Sharpe Ratio" (2014). Normal CDF via the Hart/West rational approximation; inverse via Acklam's algorithm with a Halley refinement step.* ---------------------------------------------------------------------- # Monte Carlo Equity Curve Simulator URL: https://thequant.space/tools/equity-curve-simulator/ Section: Tools Description: Simulate thousands of trading equity curves from win rate, reward-to-risk, and position size. Percentile bands, drawdown statistics, and risk of ruin — all in your browser. Two traders with the *same* strategy — same win rate, same reward-to-risk — can end the year with wildly different equity curves. This simulator makes that variance visible. Set your parameters, run a few thousand curves, and watch what luck alone does to identical edges. Everything runs in your browser; nothing is uploaded.

Parameters

Win: gain = risked × ratio. Loss: lose the risked amount. Equity floors at $0.

Equity curves — band = 5th–95th percentile, bold = median

## What the parameters teach you **Risk per trade is the volatility dial.** Rerun the default (1% risk) and then try 5% and 10%. The expectation per trade is identical, but the 5th-percentile curve — the unlucky trader with your exact strategy — collapses, and the ruin percentage climbs. This is the entire argument for small position sizes, made visually. **Win rate and reward-to-risk are one number in disguise.** A 35% win rate at 3:1 and a 65% win rate at 0.85:1 have similar expectation (+0.05R vs +0.10R per trade), but they produce very different *paths*: the low-win-rate curve endures brutal losing streaks (watch the max-consecutive-losses stat) that most traders abandon in practice. Simulate both before deciding which one you can psychologically hold. **Percent-risk cannot mathematically ruin you; fixed-risk can.** With percent sizing, losses shrink your bets — equity decays but never hits zero. With fixed-dollar sizing, a losing streak takes you to zero directly. Reality is in between: minimum position sizes, gaps, and fees make percent sizing less protective than the simulation shows. That gap between model and reality is exactly what the [production checklist](/guides/production-checklist/)'s drawdown limits exist for. **The band is the strategy; a single backtest is one thread of it.** Any individual equity curve — including your live track record — is one draw from a distribution like this fan. Judging a strategy (or a fund manager) on one path without knowing the fan's width is how survivorship stories get sold. The same logic underpins our skepticism about single-backtest papers in the [research hubs](/topics/). *Model: each trade risks the configured amount; wins pay risk × ratio, losses lose the risk; equity floors at $0. Trades are independent — real markets have regime clustering, so treat these bands as optimistic. This tool rebuilds the equity-curve simulator formerly hosted on the author's personal site (2018–2023), with percentile bands and ruin statistics added.* ---------------------------------------------------------------------- # Quant Researcher Compute-Cost Calculator URL: https://thequant.space/tools/compute-cost-calculator/ Section: Tools Description: Interactive calculator: local GPU vs cloud GPU vs API costs for quant research. Editable assumptions, break-even hours, and a monthly budget for your whole stack. Every default below is editable and dated **September 2026** — prices drift, so adjust before deciding. The reasoning behind the model is in the [local vs cloud GPU guide](/guides/local-vs-cloud-gpu/); the rest of the budget lines come from the [research stack guide](/guides/quant-research-stack/).

Your workload

Results

Own the GPU

$—

    Rent the GPU

    $—
      —
      Assumptions (edit me — defaults dated September 2026)
      ## How the model works **Own** = card price amortized linearly (default 30 months, conservative — cards hold resale value), plus metered electricity, plus local NVMe amortized over 3 years. **Rent** = marketplace-cloud hourly rate (Vast/RunPod-tier interruptible; hyperscaler on-demand runs 2–5× higher), plus cloud object storage. Both share your data subscriptions, batch-API spend, and VPS line so the totals are whole-stack budgets, not just GPU line items. What it deliberately ignores: your time (ops overhead of owning, checkpoint discipline of renting), resale value (favors owning), and data-licensing restrictions that can forbid third-party clouds outright — that last one overrides any price math and is covered in the [guide](/guides/local-vs-cloud-gpu/#the-two-constraints-generic-guides-miss). **Prefer a spreadsheet?** Download the same model as an .xlsx with live formulas — edit the blue cells, keep it in your planning folder. ---------------------------------------------------------------------- # Tick Data Storage Sizer & Cost Estimator URL: https://thequant.space/tools/tick-data-storage-sizer/ Section: Tools Description: Interactive tick data storage calculator: raw and compressed GB per day, year, and total for trades, quotes, L2 depth, or full order book across your universe — with October 2026 cost estimates for NVMe, S3, ClickHouse Cloud, Timescale Cloud, QuestDB Cloud, and a Hetzner box, plus the data-feed bill. Most tick-storage decisions go wrong before the first byte is written: either a cluster is bought for a dataset that fits on one NVMe drive, or a laptop is pointed at a full-depth feed that produces a terabyte a week. This sizer does the arithmetic. Every default is editable and dated **October 2026** — message rates and list prices both drift, so adjust before deciding. The storage layout that makes the compression ratio real is in [how to store tick data efficiently](/guides/store-tick-data-efficiently/); the engine choice is in the [tick database comparison](/guides/tick-data-databases/). Everything runs in your browser; nothing is uploaded.

      Your data

      Results

      Total compressed archive (single copy)

      —

        Per day

        —

          Per year (at today's rate)

          —
            —

            What it costs to keep — storage tiers at October 2026 list prices

            Where it livesStoredStorage $/monthFirst-year storageFirst year incl. feed

            Assumptions (edit me — defaults dated October 2026)
            Encoding and growth
            Storage prices ($/GB-month unless stated) — approximate list prices, verify before buying

            Rate defaults are order-of-magnitude, October 2026: equities use the consolidated tape for trades/quotes and a single primary venue for depth; crypto and FX are per pair on one venue. Monthly storage cost is computed at the full archive size (a backfill loaded up-front), so first-year storage = 12 × monthly. Compute for managed databases is a flat placeholder; your real bill depends on query load. Nothing here is sponsored.

            ## How the model works **Volume** = symbols × events per symbol per day × bytes per event, in an uncompressed typed record; divide by the compression ratio for what lands on disk as sorted zstd Parquet. Years are summed with the growth factor applied to each later year. The ingestion rate spreads the day's events over the trading hours; the peak multiplies that by a burst factor. **Cost** is monthly storage at the full archive size (a backfill loaded up-front) times twelve, plus the market-data feed times twelve, for an all-in first-year number. Self-managed lines (your NVMe, the Hetzner box) carry the replication factor because you are your own durability; object storage and managed databases replicate internally and are charged on a single logical copy. Managed-database compute is a flat placeholder for the smallest always-on service — real bills scale with how hard you query. **The verdict** uses two thresholds that are themselves assumptions: about 2 TB compressed and 20k events/s average for the laptop tier, about 20 TB and 200k events/s for one server. They come from what a single NVMe drive and a single process comfortably handle in 2026; argue with them if your hardware is unusual. What it deliberately ignores: derived layers (bars, features — regenerable and usually smaller than the ticks), egress fees when you pull from object storage, your own time operating a server, and the licensing layer on top of the feed price. The [storage guide](/guides/store-tick-data-efficiently/) explains why sort order and integer prices are what make the 6× compression default real; the [database comparison](/guides/tick-data-databases/) explains which engine fits which workload shape once you know the tier; [Parquet vs a database](/guides/parquet-vs-database/) covers when files stop being enough; and the [vendor guide](/guides/market-data-vendors/) covers the feed line — survivorship, point-in-time integrity, and licensing matter more than the monthly price. *Rate defaults: order-of-magnitude figures for liquid names as of October 2026; a single mega-cap or BTC/USDT alone runs 5–20× the per-symbol average. Prices: approximate public list prices, October 2026, no sponsored placements; if that ever changes it will be disclosed inline.* ---------------------------------------------------------------------- # Topic hub: Commodities & Energy Markets URL: https://thequant.space/topics/commodities-energy/ Papers: 203 (data: https://thequant.space/data/topics/commodities-energy.json) Description: Research on crude oil, natural gas, electricity and power markets, carbon pricing, metals, and agricultural commodities — pricing, forecasting, hedging, and trading. Commodity and energy markets break the assumptions that equity research takes for granted: the underlying is **physical**, storage and delivery matter, prices spike and mean-revert on seasonal and weather cycles, and in electricity the asset cannot be stored at all. That is why this literature has its own models — futures-curve dynamics, storage-constrained pricing, spike-jump processes for power, and the carbon-market mechanics of emission allowances — and why volatility and forecasting results from equities rarely transfer. The rigor split here is stark. Forecasting papers (oil, gas, power prices) often test against naïve and seasonal baselines on real data and score well. Pricing and hedging papers for structured energy contracts tend to be theory-first. When a paper claims a trading result, check whether it uses **tradable front-month contracts with roll costs** or a spliced continuous series that never existed, and whether the sample includes 2020–2022, the stress test nothing in this market escaped. Related hubs: [Volatility](/topics/volatility/), [Options & Derivatives](/topics/options-derivatives/), [Machine Learning for Forecasting](/topics/machine-learning/), [Risk Management](/topics/risk-management/). ---------------------------------------------------------------------- # Topic hub: Crypto Markets & DeFi URL: https://thequant.space/topics/crypto-defi/ Papers: 603 (data: https://thequant.space/data/topics/crypto-defi.json) Description: Research on cryptocurrency markets, DeFi protocols, AMMs, stablecoins, and blockchain market structure. Crypto gives researchers something traditional markets never could: **fully transparent, public market data** — every trade, every liquidity position, every liquidation, on-chain and downloadable. The result is a research field where microstructure questions that took decades to answer for equities get answered in months, alongside genuinely new objects of study: automated market makers, MEV, stablecoin mechanisms, and perpetual futures funding. Quality varies more here than in any other hub. The strongest work treats crypto as a **natural laboratory** — testing arbitrage, momentum, and market-making theory against complete data. The weakest imports an equity-market technique, runs it on three years of BTC prices, and declares alpha. Check the sample period against regime shifts (pre/post 2021 institutional entry, exchange collapses) before trusting any backtest, and note that many exploitable inefficiencies documented in 2018-era papers have long since been arbitraged away. Related hubs: [Market Microstructure](/topics/market-microstructure/), [Statistical Arbitrage](/topics/statistical-arbitrage/), [NLP & LLMs in Finance](/topics/nlp-llm/). ---------------------------------------------------------------------- # Topic hub: Factor Investing & Asset Pricing URL: https://thequant.space/topics/factor-investing/ Papers: 363 (data: https://thequant.space/data/topics/factor-investing.json) Description: Research on factor models, return anomalies, momentum, the cross-section of returns, and empirical asset pricing. Asset pricing asks the field's central question — **why do some assets earn higher returns than others** — and factor investing is its practical export. This hub collects the cross-sectional literature: momentum, value, carry, quality, low-volatility, the factor-zoo debates, and machine-learning approaches to the cross-section of expected returns. The elephant in this literature is **replication decay**. Documented anomalies shrink dramatically after publication, and a large fraction fail replication once you correct for data mining across hundreds of tested signals. That makes publication date and multiple-testing correction the two most informative fields in any factor paper: work that applies higher t-statistic thresholds, or tests on post-publication and international samples, deserves far more weight than a fresh anomaly with a 1990-2015 US backtest. The rigor ranking below tends to surface exactly those papers. Related hubs: [Portfolio Optimization](/topics/portfolio-optimization/), [Machine Learning for Forecasting](/topics/machine-learning/), [Fixed Income & Rates](/topics/fixed-income/). ---------------------------------------------------------------------- # Topic hub: Fixed Income & Interest Rates URL: https://thequant.space/topics/fixed-income/ Papers: 351 (data: https://thequant.space/data/topics/fixed-income.json) Description: Research on yield curves, term structure models, bond markets, and credit spreads. Fixed income is where quantitative finance's modeling tradition runs deepest and data runs thinnest. This hub covers **term-structure models** (affine, HJM-lineage, and their ML successors), **yield-curve forecasting**, bond risk premia, and **credit** — spread dynamics, default modeling, and corporate bond market structure. Two things distinguish strong papers here. First, honest treatment of the **small-sample problem**: there are only a handful of independent interest-rate cycles in any dataset, so out-of-sample claims rest on far less evidence than equity research enjoys, and papers that acknowledge regime dependence (pre/post-2008 zero-rate era, the 2022 repricing) age better. Second, attention to **market plumbing** — corporate bond research that ignores dealer intermediation costs and stale quotes produces backtests no desk can trade. The intersection with machine learning is growing fast, but the base rates favor parsimonious models in this asset class more than anywhere else. Related hubs: [Risk Management](/topics/risk-management/), [Options & Derivatives](/topics/options-derivatives/), [Factor Investing](/topics/factor-investing/). ---------------------------------------------------------------------- # Topic hub: High-Frequency Trading & Optimal Execution URL: https://thequant.space/topics/high-frequency-trading/ Papers: 337 (data: https://thequant.space/data/topics/high-frequency-trading.json) Description: Research on HFT, optimal execution, TWAP/VWAP algorithms, and intraday trading strategies. This hub spans two tightly linked literatures. **Optimal execution** — how to trade a large order while balancing market impact against timing risk — is a stochastic-control problem with a celebrated closed-form lineage from Almgren-Chriss onward, and it remains one of the most directly implementable areas of quant research. **High-frequency trading** research studies the strategies, profitability, and market effects of trading at millisecond horizons. For execution papers, the gap to watch is between model and market: elegant impact models assume permanent/temporary impact splits that are hard to estimate from data you can actually get. Papers scoring high on rigor calibrate against real parent-order data or realistic simulators rather than assuming parameters. For HFT studies, dataset provenance is everything — results from proprietary exchange data with participant identifiers say far more than inferences from public trades-and-quotes feeds. Related hubs: [Market Microstructure](/topics/market-microstructure/), [Reinforcement Learning](/topics/reinforcement-learning/), [Statistical Arbitrage](/topics/statistical-arbitrage/). ---------------------------------------------------------------------- # Topic hub: Insurance & Actuarial Risk URL: https://thequant.space/topics/insurance-actuarial/ Papers: 302 (data: https://thequant.space/data/topics/insurance-actuarial.json) Description: Research on insurance pricing, reinsurance, annuities and pensions, mortality and longevity risk, solvency, and risk sharing — scored for math complexity and empirical rigor. Insurance research shares its toolkit with trading research — risk measures, stochastic control, extreme-value statistics — but asks different questions: how to **price a liability whose payoff is a claim rather than a market quote**, how to share risk between insurer, reinsurer, and policyholder, how to fund annuities when people live longer than the tables said, and how much capital keeps the firm solvent at a regulatory confidence level. Papers here range from premium-principle theory to catastrophe-bond pricing and pension-fund asset allocation. Read with two filters. First, **which data the paper touches**: actuarial work often validates on simulated claim processes, so a paper that fits real loss or mortality data earns its rigor score. Second, **which regulatory frame it assumes** — Solvency II, Swiss Solvency Test, or risk-based capital rules change the risk measure, the horizon, and the answer. Risk-measure theory with insurance motivation also lives in the [Risk Management](/topics/risk-management/) hub; the two overlap on purpose. Related hubs: [Risk Management & Tail Risk](/topics/risk-management/), [Stochastic Control & Optimal Stopping](/topics/stochastic-control/), [Fixed Income & Interest Rates](/topics/fixed-income/). ---------------------------------------------------------------------- # Topic hub: Machine Learning for Market Forecasting URL: https://thequant.space/topics/machine-learning/ Papers: 1253 (data: https://thequant.space/data/topics/machine-learning.json) Description: Research using deep learning, LSTMs, transformers, and tree ensembles to forecast returns, prices, and market states. The largest hub in the collection, covering neural networks, tree ensembles, transformers, and hybrid models applied to **return prediction, price forecasting, and market-state classification**. The volume reflects the field's popularity; the score distribution reflects its problem — a large share of papers demonstrate in-sample pattern-fitting rather than out-of-sample edge. A practical filter when reading: the base rate for daily directional accuracy on liquid instruments hovers near 50%, so headline accuracies of 60%+ usually signal **look-ahead bias, survivorship bias, or an illiquid universe** rather than alpha. The papers ranked highest below are those that publish full train/validation/test splits, compare against naive baselines, and account for costs. Tree ensembles on engineered features remain stubbornly competitive with deep architectures on tabular financial data — a finding that replicates across this hub repeatedly. Related hubs: [Reinforcement Learning](/topics/reinforcement-learning/), [NLP & LLMs in Finance](/topics/nlp-llm/), [Factor Investing](/topics/factor-investing/). ---------------------------------------------------------------------- # Topic hub: Market Microstructure URL: https://thequant.space/topics/market-microstructure/ Papers: 579 (data: https://thequant.space/data/topics/market-microstructure.json) Description: Research on limit order books, market making, price impact, and liquidity — scored for math complexity and empirical rigor. Market microstructure studies **how prices actually form**: the dynamics of limit order books, the economics of market making, price impact of large orders, and the information content of order flow. It is the layer where trading strategy meets market plumbing, and it matters to anyone whose backtest assumes fills at mid-price. Two branches dominate this collection. The **theoretical** branch models order books with queueing theory and stochastic control — mathematically rich, but check whether the model's assumptions (Poisson arrivals, constant spreads) survive contact with data. The **empirical** branch measures impact curves, adverse selection, and liquidity patterns from tick data; its perennial weakness is dataset access, so pay attention to which venue and period a result comes from before assuming it generalizes. Related hubs: [High-Frequency Trading & Execution](/topics/high-frequency-trading/), [Statistical Arbitrage](/topics/statistical-arbitrage/), [Crypto Markets & DeFi](/topics/crypto-defi/). ---------------------------------------------------------------------- # Topic hub: Market Simulation & Agent-Based Models URL: https://thequant.space/topics/market-simulation/ Papers: 226 (data: https://thequant.space/data/topics/market-simulation.json) Description: Research on agent-based market simulators, synthetic order books and price series, generative market models, LLM trading agents, and backtesting engines. When history is one path and you need many, you simulate. This hub collects three generations of that idea: **agent-based models** (zero-intelligence and heterogeneous traders interacting through an order book), **generative models** (GANs, diffusion models, and other learned simulators of prices or order flow), and the newest wave, **LLM-driven trading agents** placed in a synthetic market to study behaviour and crowding. Alongside them sit the engineering papers on backtesting engines and market "digital twins". A simulator is only as useful as the questions it can answer, so judge each paper by what it validates. Reproducing stylised facts (fat tails, volatility clustering, the volume–volatility relation) is table stakes, not evidence that a strategy tested inside the simulator would survive real fills. Look for **calibration to real market data**, for out-of-sample tests of the simulator itself, and for honesty about market impact — the thing simulators exist to study and the thing most backtests silently ignore. Our [simulation checklist](/tags/trading-simulation/) spells this out. Related hubs: [Market Microstructure](/topics/market-microstructure/), [Reinforcement Learning for Trading](/topics/reinforcement-learning/), [NLP & LLMs in Finance](/topics/nlp-llm/), [HFT & Optimal Execution](/topics/high-frequency-trading/). ---------------------------------------------------------------------- # Topic hub: NLP & LLMs in Finance URL: https://thequant.space/topics/nlp-llm/ Papers: 509 (data: https://thequant.space/data/topics/nlp-llm.json) Description: Research on sentiment analysis, news analytics, and large language models applied to trading and financial analysis. Text became a mainstream data source for quant research in three waves: dictionary-based **sentiment scoring** of news and filings, supervised NLP on earnings calls and social media, and — since 2023 — **large language models** used for everything from extracting signals to simulating analysts. This hub tracks all three, with the LLM wave now dominating new submissions. The evaluation bar shifts with each wave, and it's worth knowing where it currently stands: for LLM-based trading papers, the critical flaw to check is **temporal leakage** — a model whose training data includes the test period's news has seen the future, and many published backtests fail exactly this test. The papers ranked highest below either use point-in-time text data with pre-cutoff models or study tasks (extraction, classification, summarization quality) where leakage doesn't invalidate the result. Sentiment-signal decay is the other recurring theme: documented text signals fade fast once published. Related hubs: [Machine Learning for Forecasting](/topics/machine-learning/), [Crypto Markets & DeFi](/topics/crypto-defi/), [Factor Investing](/topics/factor-investing/). ---------------------------------------------------------------------- # Topic hub: Options & Derivatives Pricing URL: https://thequant.space/topics/options-derivatives/ Papers: 595 (data: https://thequant.space/data/topics/options-derivatives.json) Description: Research on option pricing, hedging, implied volatility, and derivatives markets — from Black-Scholes extensions to deep hedging. Derivatives pricing is quantitative finance's oldest deep-math tradition, and this hub holds the collection's densest concentration of stochastic calculus. The research spans **pricing models** (jump diffusions, stochastic volatility, rough volatility), **numerical methods** (Monte Carlo variants, PDE solvers, ML surrogates for high-dimensional pricing), and the newer **deep hedging** literature that learns hedge ratios directly under transaction costs. The quadrant system earns its keep here: much of this work is mathematically airtight but empirically untested ("Lab Rats" in our scoring), which is fine for pricing engines and dangerous for trading claims. When a paper proposes a new model, the questions that matter are calibration stability — does it fit today's surface without absurd parameters tomorrow — and hedging performance out of sample, not merely in-sample fit to option quotes. Related hubs: [Volatility Modeling](/topics/volatility/), [Fixed Income & Rates](/topics/fixed-income/), [Machine Learning for Forecasting](/topics/machine-learning/). ---------------------------------------------------------------------- # Topic hub: Portfolio Optimization URL: https://thequant.space/topics/portfolio-optimization/ Papers: 633 (data: https://thequant.space/data/topics/portfolio-optimization.json) Description: Research on portfolio construction, mean-variance optimization, risk parity, and allocation under uncertainty — scored for rigor and complexity. Seventy years after Markowitz, portfolio optimization is still an open problem — not because the math is unsettled, but because **estimation error eats the theoretical gains**. The literature here spans classical mean-variance and its regularized descendants, risk parity and volatility targeting, Black-Litterman-style view blending, and newer machine-learning approaches to covariance estimation and end-to-end allocation. The recurring finding worth internalizing: naive equal-weight portfolios are shockingly hard to beat out of sample, so a paper's value usually lies in **how honestly it handles estimation risk** — shrinkage, robust optimization, turnover penalties — rather than in the elegance of its objective function. The ranking below favors papers that test allocations on real data with realistic constraints. Related hubs: [Factor Investing & Asset Pricing](/topics/factor-investing/), [Risk Management](/topics/risk-management/), [Machine Learning for Forecasting](/topics/machine-learning/). ---------------------------------------------------------------------- # Topic hub: Reinforcement Learning for Trading URL: https://thequant.space/topics/reinforcement-learning/ Papers: 319 (data: https://thequant.space/data/topics/reinforcement-learning.json) Description: Research applying reinforcement learning — DQN, policy gradients, actor-critic — to trading, execution, and portfolio problems. Reinforcement learning promises to learn trading policies **directly from market interaction** rather than through the forecast-then-optimize pipeline. The papers here apply DQN, policy gradients, and actor-critic methods to portfolio allocation, optimal execution, market making, and hedging. Read this literature with calibrated skepticism: markets are non-stationary, near-random-walk environments with weak reward signals — close to the hardest possible setting for RL. The papers that score well on rigor tend to share three habits: they train and test on **strictly separated periods**, they benchmark against simple baselines (buy-and-hold, equal-weight, TWAP) rather than only against other neural approaches, and they report sensitivity to transaction costs. A Sharpe ratio from an RL agent without those three checks is a screenshot, not a result. Where the field genuinely shines is **execution and hedging** — problems with dense feedback and clear cost structure. Related hubs: [Machine Learning for Forecasting](/topics/machine-learning/), [High-Frequency Trading & Execution](/topics/high-frequency-trading/), [Portfolio Optimization](/topics/portfolio-optimization/). ---------------------------------------------------------------------- # Topic hub: Risk Management & Tail Risk URL: https://thequant.space/topics/risk-management/ Papers: 823 (data: https://thequant.space/data/topics/risk-management.json) Description: Research on VaR, expected shortfall, tail risk, systemic risk, and stress testing — scored for empirical rigor. Risk management research divides into measurement and mechanism. The **measurement** literature — VaR and expected shortfall estimation, extreme value theory, drawdown analysis — asks how to quantify what can go wrong in a portfolio. The **mechanism** literature — systemic risk, contagion networks, stress testing — asks how losses propagate between institutions and markets. The practical test for measurement papers is backtesting honesty: a VaR model is only as good as its **violation rate and the independence of its violations**, and papers that report Kupiec or Christoffersen tests on out-of-sample data outrank those that stop at in-sample fit. For tail-risk work specifically, be wary of estimators tuned to one crisis; the ranking below rewards papers validated across multiple stress regimes. This is also where regulation (Basel's shift from VaR to expected shortfall) keeps the academic work unusually connected to practice. Related hubs: [Volatility Modeling](/topics/volatility/), [Portfolio Optimization](/topics/portfolio-optimization/), [Fixed Income & Rates](/topics/fixed-income/). ---------------------------------------------------------------------- # Topic hub: Statistical Arbitrage & Pairs Trading URL: https://thequant.space/topics/statistical-arbitrage/ Papers: 89 (data: https://thequant.space/data/topics/statistical-arbitrage.json) Description: Research on statistical arbitrage, pairs trading, cointegration, and mean-reversion strategies — scored for math complexity and empirical rigor. Statistical arbitrage covers strategies that trade **temporary mispricings between related instruments**: classic pairs trading, cointegration-based baskets, and mean-reversion signals across equities, futures, and crypto. It is one of the few strategy families where academic papers regularly include real backtests, which makes the rigor score especially useful here. When reading this literature, three questions separate implementable work from curve-fitting: Does the paper account for **transaction costs and shorting constraints**? Is the pair-selection rule defined **before** the backtest window? And does performance survive the post-2010 period, when the classic distance method famously decayed? Papers ranked highly below tend to answer at least two of these. Related hubs: [Market Microstructure](/topics/market-microstructure/), [Machine Learning for Forecasting](/topics/machine-learning/), [High-Frequency Trading](/topics/high-frequency-trading/). ---------------------------------------------------------------------- # Topic hub: Stochastic Control & Optimal Stopping URL: https://thequant.space/topics/stochastic-control/ Papers: 512 (data: https://thequant.space/data/topics/stochastic-control.json) Description: Research on stochastic control, optimal stopping, HJB equations, mean-field games, and BSDEs applied to investment, consumption, liquidation, and dividend problems. This hub holds the mathematical engine room of quantitative finance: **stochastic control** (choose a policy that steers a diffusion), **optimal stopping** (choose a time to act), and the machinery that solves them — Hamilton–Jacobi–Bellman equations, viscosity solutions, backward stochastic differential equations, and mean-field games for the many-player limit. The classic applications are Merton-style consumption and investment, optimal liquidation, dividend and reinsurance control, and American-option exercise. Most papers here are **Lab Rats** by our scoring: deep theory, little or no data. That is not a criticism of the work, but it changes how you read it. The questions that matter are whether the model's state variables are *observable* in practice, whether the closed-form or numerical solution degrades gracefully when parameters are estimated rather than known, and whether a paper that claims a strategy ever confronts it with transaction costs or discrete rebalancing. The handful of papers that pair a control result with a calibrated numerical study rank highest below. Related hubs: [Options & Derivatives](/topics/options-derivatives/), [Portfolio Optimization](/topics/portfolio-optimization/), [HFT & Optimal Execution](/topics/high-frequency-trading/), [Insurance & Actuarial Risk](/topics/insurance-actuarial/). ---------------------------------------------------------------------- # Topic hub: Volatility Modeling & Forecasting URL: https://thequant.space/topics/volatility/ Papers: 746 (data: https://thequant.space/data/topics/volatility.json) Description: Research on GARCH, realized volatility, rough volatility, the VIX, and volatility forecasting across asset classes. Volatility is the most forecastable quantity in finance — unlike returns, it clusters, mean-reverts, and leaves a measurable footprint in high-frequency data. This hub collects the modeling lineage from **GARCH** through **realized volatility** built on intraday data, to **rough volatility** and the modern deep-learning forecasters, along with VIX and variance-risk-premium research. Because volatility forecasting has real benchmarks, rigor is easier to judge here than elsewhere: a serious paper compares against HAR-RV (the stubborn workhorse baseline) on out-of-sample data, not just against its own ablations. Watch also for the **evaluation-loss trap** — rankings of volatility models flip depending on whether you score them with MSE, QLIKE, or economic value in an options or risk-targeting strategy, and good papers say so explicitly. Related hubs: [Options & Derivatives](/topics/options-derivatives/), [Risk Management](/topics/risk-management/), [Machine Learning for Forecasting](/topics/machine-learning/).