Paper: arXiv 2610.07888
Authors: Jonas Brunholm, Bjarne Højgaard, Thomas D. Nielsen, Orimar Sauri
Abstract
This paper proposes a framework for modeling probability of default via functional data analysis. By representing a series of credit variables as functions, we investigate whether intra-monthly information improves default predictability in linear models. We further show how a range of widely used variables, among them available funds, utilization rate, and overdraft, can all be derived from three quantities observed over time. Namely, 1) the type of each account, 2) the balance of the account, and 3) the size of the credit limit on the account. Retaining these quantities as continuous-time processes, rather than reducing them to monthly aggregated values, yields a continuous faithful representation of the borrower. When analyzing these three processes, we discovered that recurring events associated with the ordinal position among banking days created strong cyclical patterns. We therefore develop a relative time framework that aligns the recurring events across borrowers, ensuring that borrowers possess the same cyclical pattern, regardless of real time. We assess the framework using functional logistic regression. This approach accommodates the continuous representation while remaining closely related to a logistic regression model commonly used in credit risk practice, due to strict regulatory constraints. We show that, when equipped with an effective functional representation of transactional trajectories, the proposed model attains predictive performance on par with XGBoost while consistently outperforming logistic regression. Importantly, the model balances predictive performance of a machine learning model with the interpretability of linear default models, potentially enabling financial institutions to use the model, even under strict regulation.
Complexity vs Empirical Score
- Math Complexity: 7.0/10
- Empirical Rigor: 8.0/10
- Quadrant: Holy Grail — high math complexity, high empirical rigor
Why this score: This paper presents a novel approach to PD modeling using functional data analysis, demonstrating strong empirical results. The methodology is well-explained and addresses a practical problem in credit risk with an interpretable solution. The combination of advanced mathematical techniques with robust empirical validation places it in the ‘Holy Grail’ quadrant.
Research Flowchart
flowchart TD
A[Research Goal: Improve PD Modeling with Intra-Monthly Info & Functional Data] --> B(Methodology: Functional Data Analysis)
B --> C{Inputs: Credit Account Type, Balance, Credit Limit (Continuous Time)}
C --> D[Computational Process: Functional Logistic Regression & Relative Time Framework]
D --> E{Outcomes: Improved Predictability, Cyclical Pattern Alignment}
E --> F[Key Findings: Performance on par with XGBoost, Outperforms Logistic Regression, Interpretable]