Paper: arXiv 2609.27654
Authors: Sultan Amed, Tanmay Sen, Sayantan Banerjee
Abstract
Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.
Complexity vs Empirical Score
- Math Complexity: 6.0/10
- Empirical Rigor: 8.0/10
- Quadrant: Holy Grail — high math complexity, high empirical rigor
Why this score: This paper presents a robust empirical study of federated learning for income estimation, a novel application in digital lending. The methodology is clearly explained, and the use of a large, partitioned dataset for simulation enhances its empirical rigor. While the mathematical derivations are not explicitly detailed in the summary, the underlying federated learning framework and statistical analysis suggest a solid mathematical foundation.
Research Flowchart
flowchart TD
A[Research Goal: Estimate Income in Digital Lending Under Data Sovereignty Constraints] --> B{Key Methodology: Federated Learning (FedIncome) Framework};
B --> C[Data/Inputs: LendingClub Loans (1M+), Partitioned into 50 State-Level Clients];
C --> D{Computational Process: Train Shared Model without Pooling Raw Data, Simulate Heterogeneous Consortium};
D --> E[Key Finding 1: Best FedIncome model achieves Out-of-Time R^2 = 0.608 (vs. 0.619 for pooled benchmark)];
E --> F[Key Finding 2: Small-sample clients gain 3.8 ppt R^2 improvement; Crossover at ~4,790 training observations];
F --> G[Key Finding 3: Federation increases simulated approval rates with modest changes in default rates (using FedIncome estimates + DTI thresholds)];