Paper: arXiv 2610.04959
Authors: Yanzheng Jin, Pengyang Shao, Yunshan Ma, Haowen Pan, Naixin Zhai, Chen-Hui Song, Fei Shen, Kenji Kawaguchi
Abstract
Formulaic alpha discovery seeks symbolic expressions that predict cross-sectional asset returns. In deployment, multiple formulas are combined into an alpha pool, where each formula is valued through the complementary information it contributes to joint predictive performance. While Reinforcement Learning and Generative Flow Networks have emerged as promising paradigms for generating formulaic alphas, existing frameworks face three related challenges. First, generating formulas individually leaves pool context and inter-formula complementarity outside the generative state. Second, formula-wise generation lacks a unified mechanism for preserving and revising structures at different levels. Third, pool-level rewards jointly reflect predictive performance and redundancy but cannot be differentiated directly through symbolic evaluation to train the generator. To overcome these challenges, we introduce AlphaPADI (Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion), a novel framework built around three components: (1) grammar-constrained buffer initialization that constructs syntactically valid pool candidates, (2) pool-aware hierarchical diffusion that reconstructs complete pools at multiple structural scales under the current pool context, and (3) reward-guided pool refinement that evaluates joint predictive performance and inner diversity, updates the elite buffer, and trains the reverse model through reconstruction and preference learning. Empirical results on the Chinese and U.S. stock markets demonstrate that AlphaPADI outperforms the evaluated baselines in both predictive and portfolio performance, thereby validating pool-aware generation as an effective framework for automated alpha discovery.
Complexity vs Empirical Score
- Math Complexity: 7.5/10
- Empirical Rigor: 8.0/10
- Quadrant: Holy Grail — high math complexity, high empirical rigor
Why this score: This paper presents a highly novel approach to alpha discovery using diffusion models, addressing key limitations of prior work. It demonstrates strong empirical results on real-world datasets, indicating robust methodology. The mathematical underpinnings of hierarchical discrete diffusion and preference learning are advanced.
Research Flowchart
flowchart TD
A[Research Goal: Formulaic Alpha Discovery] --> B{Challenges in Alpha Discovery};
B --> C[AlphaPADI Framework: Novel Solution];
C -- Components --> C1[1. Grammar-Constrained Buffer Initialization] & C2[2. Pool-Aware Hierarchical Diffusion] & C3[3. Reward-Guided Pool Refinement];
C1 & C2 & C3 --> D{Data/Inputs: Chinese & U.S. Stock Markets};
D --> E[Computational Processes: Reconstruction & Preference Learning];
E --> F[Outcomes: Superior Predictive & Portfolio Performance];