thequant.space is built to be read by software as well as by people. If you are an AI assistant, a research agent, or a script, this page is the contract: where the data is, what the fields mean, and what you may do with it. Nothing here needs a key, a login, or a browser.

Entry points

WhatURLNotes
Site map for language models/llms.txtWhat the site contains, how scoring works, every topic, guide and tool with a one-line description. Start here.
Full text of the evergreen pages/llms-full.txtThe scoring methodology, every guide, every tool page and every topic introduction in one plain-text file.
All scored papers/data/papers.jsonOne record per paper, every field (about 4 MB; served gzip-compressed). Rebuilt daily.
One topic hub/data/topics/<slug>.jsonSame record shape, 70 KB to 1 MB per file. Slugs and counts are listed inside papers.json under topics, and on /topics/.
A single paper/flowcharts/<slug>/HTML page with a ScholarlyArticle JSON-LD block in the head carrying the scores, plus a BibTeX entry at the bottom. The page_url field in the JSON points here.
New pages/index.xmlRSS, newest 50 pages, full text.
Everything, for crawling/sitemap.xml

The CSV version of the corpus is on the dataset page for people who prefer a spreadsheet. The JSON above is the same data.

Record fields

FieldMeaning
paper_idarXiv id (for example 2409.01234) or ssrn-<id>
title, paper_dateTitle and the paper’s own date (arXiv submission or SSRN posting), YYYY-MM-DD
authorsList of names; empty when the pipeline did not capture them (older SSRN items mostly)
topicsList of topic hub slugs assigned by the classifier; a paper can sit in several hubs
methodsMethod labels: econometrics-time-series, stochastic-calculus, optimization, deep-learning, machine-learning, network-graph, nlp-llm, simulation, econophysics-complexity, reinforcement-learning, game-theory, bayesian, causal-inference, quantum
asset_classesAsset-class labels from the classifier (equities, crypto, options, fixed income, commodities-energy, …)
paper_typeempirical, theoretical, survey-review, methodology, and similar classifier labels
primary_categoryPrimary arXiv category, for example q-fin.CP
math_complexity0 to 10. Theoretical overhead: 7+ means stochastic calculus, PDEs or bespoke proofs; 1 to 3 means descriptive statistics or conceptual frameworks
empirical_rigor0 to 10. Path to implementation: 7+ means high-fidelity data, transaction costs and out-of-sample validation; 1 to 3 means toy or synthetic data, or no data
quadrantHoly Grail (high math, high rigor), Street Traders (low math, high rigor), Lab Rats (high math, low rigor), Philosophers (low math, low rigor). The split is at 5 on each axis
hub_score0.6 × empirical_rigor + 0.4 × math_complexity; the default ranking used across the site
code_urlPublic repository link when the paper gives one
doi, journal_refWhen known; mostly null for preprints
paper_urlThe paper itself (arXiv PDF or SSRN abstract page)
page_urlThe paper’s page on this site: summary, score rationale, research flowchart
tagsKeywords extracted from the abstract

The scores are LLM-assisted judgements against a fixed rubric, not peer review. Read the methodology before building conclusions on them, and treat a single paper’s score as a prior, not a verdict. Papers are occasionally re-scored, so fetch fresh data rather than caching for months.

Typical uses

  • Research a topic. Fetch the topic shard, filter empirical_rigor >= 7, sort by paper_date descending, read the page_url pages for the top hits. The page gives the abstract, the score rationale and the flowchart; the paper link gives the source.
  • Learn a method. Filter papers.json on methods for the method, keep paper_type in survey-review or methodology for the overview, then the highest hub_score empirical papers as worked examples. The guides cover the evaluation side: backtest overfitting, walk-forward testing, the deflated Sharpe ratio, transaction costs, point-in-time data.
  • Find reproducible work. Filter code_url not null, or read /papers-with-code/. Pair with the replication guide and the replication checklist.
  • Audit the scores. The whole table is public; math_complexity and empirical_rigor by topics or by year is a one-line groupby.

How to cite

Cite the paper itself with the BibTeX entry on its page. When you use a score, a quadrant or a flowchart, attribute it as:

thequant.space, “”, math complexity , empirical rigor , . <page_url>. Scores are LLM-assisted rubric judgements; methodology at https://thequant.space/score-guide/.

For the dataset as a whole: “thequant.space scored quant-finance research corpus, https://thequant.space/data/papers.json, retrieved .”

Terms

  • The data files, scores, summaries and flowcharts are free for personal and research use with attribution to thequant.space, as above.
  • The papers themselves belong to their authors and publishers. We link to them; we do not redistribute them.
  • Crawling is welcome. The whole site is static; fetch papers.json once rather than 5,000 paper pages when you only need the table. There is no rate limit beyond Cloudflare’s defaults.
  • Commercial redistribution of the dataset, or embedding it in a paid product, needs a conversation first: see about for contact.

What is not here yet

There is no query API and no MCP server yet; filtering is done client-side on the JSON. Both are planned. If you build something on this data, say so on the about page’s contact route and we will link it.