Who this is for. Quants moving from daily-bar research to intraday data, and anyone evaluating a microstructure or HFT paper. Assumes comfort with time-series data at scale.
How to use it. Read the steps in order; each one assumes the previous. The “why here” line says what the step adds. Where a step is a tool, run it on your own numbers before moving on. The exemplar papers at the end are the archive’s highest-rigor papers for the path’s methods; read two of them with the checklists in hand.
Steps
- OHLCV, Trade, Quote, and Order-Book Data: What Each Can Answer
The market-data hierarchy from daily bars to full order books: what each granularity can and cannot answer, the storage and cost jumps between levels, and matching data type to research question.
Why here: What OHLCV, trades, quotes and order books can and cannot tell you.
- How to Evaluate a Market-Microstructure Paper
A checklist for evaluating market-microstructure papers: dataset provenance, venue and period specificity, theoretical assumption audits, and the generalization trap.
Why here: The checklist for the literature.
- Limit-Order-Book Imbalance: What It Measures and What It Misses
Order-book imbalance as a predictor: why it works at tick horizons, the spoofing and iceberg problems, the monetization gap, and how to evaluate imbalance-based research.
Why here: The canonical short-horizon signal and its limits.
- A Visual Map of Limit-Order-Book Prediction Research
A Visual Map of Limit-Order-Book Prediction Research: the field organized into its major branches, with the strongest papers in each ranked by empirical rigor. Auto-refreshed as new papers are scored.
Why here: LOB-prediction research mapped by score.
- Market Making: Inventory Risk, Adverse Selection, and Spread Capture
The three forces of market making — spread capture, inventory risk, and adverse selection — with the Avellaneda-Stoikov intuition, the profitability identity, and the production failure catalog.
Why here: Inventory, adverse selection, spread capture.
- Transaction Costs, Slippage, and Market Impact: From Paper Alpha to Tradeable Alpha
How to model transaction costs in backtests: the four cost components, the square-root impact law, honest cost ranges by asset class, and the turnover arithmetic that kills most published strategies.
Why here: Impact models and the square-root law.
- Microstructure Noise and Realized Volatility
Why high-frequency volatility estimates explode: bid-ask bounce, discreteness, the signature plot, noise-robust estimators, and the practical sampling rules for realized volatility.
Why here: Why high-frequency volatility estimates need care.
- How to Store Tick Data Efficiently
Practical tick-data storage: partitioning schemes, columnar formats, compression choices, type discipline, and the layout decisions that make years of ticks queryable on one machine.
Why here: Storage layout for the data you now need.
- TimescaleDB vs ClickHouse vs DuckDB vs kdb+ for Tick Data Research (2026)
An honest comparison of TimescaleDB, ClickHouse, DuckDB, QuestDB, and kdb+ for storing and querying tick data in quant research — by workload, not by benchmark marketing.
Why here: Choosing the database.
- Tick Data Storage Sizer & Cost Estimator tool
Interactive tick data storage calculator: raw and compressed GB per day, year, and total for trades, quotes, L2 depth, or full order book across your universe — with October 2026 cost estimates for NVMe, S3, ClickHouse Cloud, Timescale Cloud, QuestDB Cloud, and a Hetzner box, plus the data-feed bill.
Why here: Size the storage before buying it.
- Free vs Paid Market Data: What Changes in Real Research
What separates free market data from paid: the six quality dimensions that matter for research, where free data is genuinely sufficient, and the failure modes that only surface after you've built on it.
Why here: What changes when you pay for data.
Related
All paths: learning paths. Methods, with their own guides and exemplars: methods pages. Programmatic access to the papers: agents page.