Paper: arXiv 2610.08869
Authors: Yew Lee Tan
Abstract
Regulated credit scoring requires scores monotone non-decreasing in every exposure input. Deployed pipelines – hand-crafted monotone aggregates feeding sign-constrained gradient boosting – already meet this by composition; the open question is what learned temporal aggregation is worth inside one. We answer on five production-scale credit datasets at matched admissibility (one priced baseline convention excepted), with a monotone recurrent architecture whose per-input guarantee we extend, with proofs, to vector-valued inputs and to exogenously macro-conditioned decay gates, severities, thresholds, and peak memory. Two findings result. First, a strictness ladder: the value of learned monotone features rises with governance-frame strictness – zero on unconstrained engineered panels, maximal in summaries-only frames – replicated across two datasets and an official temporal-stability metric, though unconditioned features degrade on externally adjudicated later weeks. Second, a conditioning-delivery asymmetry under regime shift. On a train-on-boom, test-on-crisis mortgage design, two public macroeconomic series hurt as input columns, yet conditioning the recurrence on them delivers the paper’s only learned-block crisis-cohort uplifts. The confirmed effect: +0.006 to +0.013 AUC on an internally pre-registered Freddie Mac replication, at all five held-out seeds. The discovery estimate: +0.015 to +0.021 on Fannie Mae (three of five seeds post hoc), worth 10-27 basis points of defaulted balance at an 80% approval cutoff, and grows with early-prepaid loans excluded. A state-level test identifies the mechanism: between-cohort calibration transfer. A pandemic-band episode bounds scope: under forbearance-distorted labels the gain generalizes at a quarter to a third of crisis size on Fannie Mae, on Freddie Mac only against the capacity control.
Complexity vs Empirical Score
- Math Complexity: 7.0/10
- Empirical Rigor: 9.0/10
- Quadrant: Holy Grail — high math complexity, high empirical rigor
Why this score: This paper presents a novel approach to credit scoring under strict regulatory constraints, extending a monotone recurrent architecture with proofs. It demonstrates high empirical rigor through extensive testing on production-scale datasets and robust statistical methodologies, including pre-registered experiments and replications. The findings offer significant practical implications for regulated financial models.
Research Flowchart
flowchart TD
A[Research Goal: Value of Learned Monotone Temporal Aggregation in Governed Credit Scoring] --> B{Key Methodology: Monotone Recurrent Architecture Extension};
B --> C[Data/Inputs: Five Production-Scale Credit Datasets, Macroeconomic Series];
C --> D{Computational Processes: Matched Admissibility, Train/Test Regimes, State-Level Tests};
D --> E[Key Finding 1: Value of Learned Features Rises with Governance-Frame Strictness];
D --> F[Key Finding 2: Conditioning-Delivery Asymmetry & Crisis-Cohort Uplifts (+0.006 to +0.013 AUC)];
F --> G[Mechanism: Between-Cohort Calibration Transfer; Pandemic-Band Scope Bounding];