Paper: arXiv 2609.34169

Authors: Kevin Foley, Jonathan Hartadi, Shivesh Prakash, Swapnil Vatsal

Abstract

News may reveal systematic risk, but whether its context enhances the construction of systematic risk factors is still unclear. We seek to test whether utilizing a sentence transformer represents an improvement over techniques such as Latent Dirichlet Allocation (LDA) in the coherence of topic term lists generated from unstructured text data. To test this, the same collection of unstructured text data comprising of 394,661 articles and the same downstream financial portfolio construction pipeline were applied with the text layer differing, including the length of article text each model used and how topic terms were ranked: we benchmark LDA against a frozen sentence transformer with k-means clustering. We find that the sentence transformer branch had higher observed scores both in terms of coherence (measured by NPMI) as well as financial performance (measured by Sharpe), although the available tests do not establish outperformance. Further exploratory specifications such as utilizing spherical clustering and multi-horizon exposures had an observed excess-return Sharpe of 1.03 for the combined model. We believe that there is some promise in applying context-aware techniques on unstructured news text, but stricter tests using only information available at each date and broader datasets may be required to enhance the confidence in the observed performance.

Complexity vs Empirical Score

  • Math Complexity: 6.5/10
  • Empirical Rigor: 7.0/10
  • Quadrant: Holy Grail — high math complexity, high empirical rigor

Why this score: This paper presents a novel comparison of text representation techniques for asset pricing, demonstrating strong empirical rigor through a well-defined pipeline and backtesting. While the mathematical derivations are not explicitly detailed in the excerpt, the underlying models (LDA, sentence transformers, Sparse IPCA) imply a significant level of mathematical sophistication. The findings, though not establishing outperformance definitively, suggest promising avenues for context-aware NLP in finance.

Research Flowchart

  flowchart TD
    A[Research Goal: Context vs. Word Counts in Topic Models for Asset Pricing] --> B(Methodology: Benchmark LDA vs. Sentence Transformer + K-Means)
    B --> C{Data: 394,661 News Articles}
    C --> D[Computational Processes: Topic Modeling, Portfolio Construction, Performance Metrics]
    D --> E{Outcomes: Higher Coherence (NPMI) & Financial Performance (Sharpe) for Sentence Transformer}
    E --> F[Conclusion: Promise in Context-Aware Techniques, but Stricter Tests Needed]