Paper: arXiv 2403.00273
Abstract
The Argentinian real estate market presents a unique case study characterized by its unstable and rapidly shifting macroeconomic circumstances over the past decades. Despite the existence of a few datasets for price prediction, there is a lack of mixed modality datasets specifically focused on Argentina. In this paper, the first edition of ARED is introduced. A comprehensive real estate price prediction dataset series, designed for the Argentinian market. This edition contains information solely for Jan-Feb 2024. It was found that despite the short time range captured by this zeroth edition (44 days), time dependent phenomena has been occurring mostly on a market level (market as a whole). Nevertheless future editions of this dataset, will most likely contain historical data. Each listing in ARED comprises descriptive features, and variable-length sets of images.
Complexity vs Empirical Score
- Math Complexity: 0.5/10
- Empirical Rigor: 8.0/10
- Quadrant: Street Traders — practical and empirical, lighter on theory
Why this score: The paper is a dataset publication with no mathematical models or derivations, but it includes detailed data scraping methodology, visualization of statistical metrics (Wasserstein distance, quantile ranges), and a plan for future updates, making it highly implementation-heavy and ready for backtesting.
Research Flowchart
flowchart TD A["Research Goal:<br>Argentine Real Estate Dataset Creation"] --> B["Methodology<br>Data Collection Jan-Feb 2024"] B --> C["Input Data<br>Property Listings + Images"] C --> D["Computational Process<br>Mixed Modality Data Aggregation"] D --> E["Key Finding 1<br>Market-Level Time-Dependent Phenomena"] D --> F["Key Finding 2<br>Lack of Historical Data"] E --> G["Outcome<br>ARED: First Edition Dataset"] F --> G