thequant.space is built to be read by software as well as by people. If you are an AI assistant, a research agent, or a script, this page is the contract: where the data is, what the fields mean, and what you may do with it. Nothing here needs a key, a login, or a browser.
Entry points
| What | URL | Notes |
|---|---|---|
| Site map for language models | /llms.txt | What the site contains, how scoring works, every topic, guide and tool with a one-line description. Start here. |
| Full text of the evergreen pages | /llms-full.txt | The scoring methodology, every guide, every tool page and every topic introduction in one plain-text file. |
| All scored papers | /data/papers.json | One record per paper, every field (about 4 MB; served gzip-compressed). Rebuilt daily. |
| One topic hub | /data/topics/<slug>.json | Same record shape, 70 KB to 1 MB per file. Slugs and counts are listed inside papers.json under topics, and on /topics/. |
| A single paper | /flowcharts/<slug>/ | HTML page with a ScholarlyArticle JSON-LD block in the head carrying the scores, plus a BibTeX entry at the bottom. The page_url field in the JSON points here. |
| New pages | /index.xml | RSS, newest 50 pages, full text. |
| Everything, for crawling | /sitemap.xml |
The CSV version of the corpus is on the dataset page for people who prefer a spreadsheet. The JSON above is the same data.
Record fields
| Field | Meaning |
|---|---|
paper_id | arXiv id (for example 2409.01234) or ssrn-<id> |
title, paper_date | Title and the paper’s own date (arXiv submission or SSRN posting), YYYY-MM-DD |
authors | List of names; empty when the pipeline did not capture them (older SSRN items mostly) |
topics | List of topic hub slugs assigned by the classifier; a paper can sit in several hubs |
methods | Method labels: econometrics-time-series, stochastic-calculus, optimization, deep-learning, machine-learning, network-graph, nlp-llm, simulation, econophysics-complexity, reinforcement-learning, game-theory, bayesian, causal-inference, quantum |
asset_classes | Asset-class labels from the classifier (equities, crypto, options, fixed income, commodities-energy, …) |
paper_type | empirical, theoretical, survey-review, methodology, and similar classifier labels |
primary_category | Primary arXiv category, for example q-fin.CP |
math_complexity | 0 to 10. Theoretical overhead: 7+ means stochastic calculus, PDEs or bespoke proofs; 1 to 3 means descriptive statistics or conceptual frameworks |
empirical_rigor | 0 to 10. Path to implementation: 7+ means high-fidelity data, transaction costs and out-of-sample validation; 1 to 3 means toy or synthetic data, or no data |
quadrant | Holy Grail (high math, high rigor), Street Traders (low math, high rigor), Lab Rats (high math, low rigor), Philosophers (low math, low rigor). The split is at 5 on each axis |
hub_score | 0.6 × empirical_rigor + 0.4 × math_complexity; the default ranking used across the site |
code_url | Public repository link when the paper gives one |
doi, journal_ref | When known; mostly null for preprints |
paper_url | The paper itself (arXiv PDF or SSRN abstract page) |
page_url | The paper’s page on this site: summary, score rationale, research flowchart |
tags | Keywords extracted from the abstract |
The scores are LLM-assisted judgements against a fixed rubric, not peer review. Read the methodology before building conclusions on them, and treat a single paper’s score as a prior, not a verdict. Papers are occasionally re-scored, so fetch fresh data rather than caching for months.
Typical uses
- Research a topic. Fetch the topic shard, filter
empirical_rigor >= 7, sort bypaper_datedescending, read thepage_urlpages for the top hits. The page gives the abstract, the score rationale and the flowchart; the paper link gives the source. - Learn a method. Filter
papers.jsononmethodsfor the method, keeppaper_typeinsurvey-reviewormethodologyfor the overview, then the highesthub_scoreempirical papers as worked examples. The guides cover the evaluation side: backtest overfitting, walk-forward testing, the deflated Sharpe ratio, transaction costs, point-in-time data. - Find reproducible work. Filter
code_urlnot null, or read /papers-with-code/. Pair with the replication guide and the replication checklist. - Audit the scores. The whole table is public;
math_complexityandempirical_rigorbytopicsor by year is a one-line groupby.
How to cite
Cite the paper itself with the BibTeX entry on its page. When you use a score, a quadrant or a flowchart, attribute it as:
thequant.space, “
”, math complexity , empirical rigor , . <page_url>. Scores are LLM-assisted rubric judgements; methodology at https://thequant.space/score-guide/.
For the dataset as a whole: “thequant.space scored quant-finance research corpus, https://thequant.space/data/papers.json, retrieved
Terms
- The data files, scores, summaries and flowcharts are free for personal and research use with attribution to thequant.space, as above.
- The papers themselves belong to their authors and publishers. We link to them; we do not redistribute them.
- Crawling is welcome. The whole site is static; fetch
papers.jsononce rather than 5,000 paper pages when you only need the table. There is no rate limit beyond Cloudflare’s defaults. - Commercial redistribution of the dataset, or embedding it in a paid product, needs a conversation first: see about for contact.
What is not here yet
There is no query API and no MCP server yet; filtering is done client-side on the JSON. Both are planned. If you build something on this data, say so on the about page’s contact route and we will link it.