Paper: arXiv 2610.03161
Authors: Nirvik Sahoo, Paul Robert Griffin
Abstract
Feature selection for imbalanced classification tasks such as credit card fraud and consumer default detection requires balancing predictive relevance, inter-feature redundancy, and computational feasibility. We benchmark three computing paradigms, classical branch-and-bound optimization (Gurobi), photonic entropy computing (QCI Dirac-3), and simulated photonic boson sampling (Piquasso), across thirteen feature-selection methods on two datasets: ULB Credit Card Fraud (30 features) and AmEx consumer default (159 features). Each method is routed to the solver matched to its mathematical structure. On ULB, Dirac-3 MI-Spearman matches the all-features model using 13 of 30 features (mean F1 0.873 +/- 0.023 over five runs, best run 0.896), and Piquasso is the best method at k=5. On AmEx, performance rises steadily with the feature budget and every paradigm approaches F1 = 0.80 only near the full feature set. Most differences between Gurobi and Dirac-3 on identical methods fall within run-to-run variation; the large gaps occur where the certified optimum generalizes poorly, most sharply for distance correlation on AmEx at k=25 (Gurobi F1 = 0.422 vs. a Dirac-3 mean of 0.746). At matched budgets, F1 varies about ten times more across methods on ULB than on AmEx, which we trace to how concentrated the predictive signal is in each feature space.
Complexity vs Empirical Score
- Math Complexity: 7.5/10
- Empirical Rigor: 8.0/10
- Quadrant: Holy Grail — high math complexity, high empirical rigor
Why this score: This paper presents a highly novel approach to feature selection using quantum solvers, backed by a robust empirical comparison across multiple methods and datasets. The mathematical underpinnings of QUBO and quantum computing are inherently complex, and the empirical setup is thorough. While code/data availability isn’t explicitly stated, the detailed methodology suggests good reproducibility.
Research Flowchart
flowchart TD
A[Research Goal: Benchmark Photonic Quantum Solvers for Feature Selection in Financial Risk Detection] --> B{Key Methodology Steps};
B --> C[Data/Inputs: ULB Credit Card Fraud (30 features) & AmEx Consumer Default (159 features)];
C --> D{Computational Processes:};
D --> D1[13 Feature Selection Methods];
D --> D2[3 Computing Paradigms: Gurobi (Classical), QCI Dirac-3 (Photonic), Piquasso (Simulated Photonic)];
D --> D3[Solvers Matched to Method Structure];
D --> D4[Performance Metric: F1 Score];
D --> E[Key Findings/Outcomes];
E --> F1[ULB: Dirac-3 MI-Spearman matches all-features model (13/30 features, F1=0.873). Piquasso best at k=5.];
E --> F2[AmEx: Performance rises with features, F1~0.80 only near full set.];
E --> F3[Gurobi vs. Dirac-3: Similar results often within run-to-run variation. Large gaps when Gurobi optimum generalizes poorly (e.g., AmEx, k=25: Gurobi F1=0.422 vs. Dirac-3 F1=0.746).];
E --> F4[F1 variation: ULB has ~10x more variation across methods than AmEx, due to predictive signal concentration.];