Paper: arXiv 2610.04040

Authors: Mingyang, Chen, Yida, Xu, Huiwen, Chen, Yiming Lu, Wei Jin

Abstract

Financial LLM agents are often evaluated by comparing their end-to-end returns with those of a baseline and testing the paired difference against zero. This measures whether deploying the agent changes realized performance, but it does not isolate event-selection skill. An agent that frequently changes positions from flat to long can earn a positive paired return from an upward-drifting event pool even when it selects events at random. We propose the Agent Policy-Value Audit, which holds fixed the observed count of each ordered action-change type and randomly reassigns them across eligible events. The average payoff from these reassignments is the composition benchmark; the difference between observed deployment value and this benchmark is selection value. In semi-synthetic benchmarks based on real earnings-event returns, a zero-centered paired test falsely attributes passive exposure to selection skill in $11.6%$ of no-skill replications, while the transition-matched audit reduces this rate to $5.3%$. Applied retrospectively to 723 earnings events at 44 U.S. consumer-facing firms, the audit decomposes the agent’s gross deployment value of $+15.2$ bps/event into a $+25.8$ composition benchmark and a $-10.6$ selection value. The agent does not detectably outperform matched random assignments. Financial-agent evaluations should report deployment value separately from event-selection value.

Complexity vs Empirical Score

  • Math Complexity: 6.0/10
  • Empirical Rigor: 8.0/10
  • Quadrant: Holy Grail — high math complexity, high empirical rigor

Why this score: This paper introduces a novel and robust methodology for evaluating financial LLM agents, addressing a critical flaw in current evaluation practices. Its strong empirical validation and clear explanation of a complex problem contribute to its high overall score.

Research Flowchart

  flowchart TD
    A[Research Goal: Isolate Event-Selection Skill in Financial LLM Agents] --> B{Traditional Evaluation Limitations};
    B -- Problem: Confounds Passive Exposure with Skill --> C[Proposed Methodology: Agent Policy-Value Audit];
    C -- Steps --> D{1. Fix observed action-change types<br>2. Randomly reassign them to eligible events<br>3. Calculate composition benchmark (average payoff)};
    D -- Inputs --> E[Real Earnings Event Returns<br>Observed Agent Actions];
    E -- Process --> F{1. Semi-synthetic benchmarks<br>2. Retrospective analysis (723 earnings events, 44 firms)};
    F -- Outcomes --> G[Key Findings:
    - Traditional evaluation: 11.6% false positives (no skill)
    - Audit: 5.3% false positives
    - Agent's Gross Deployment Value: +15.2 bps/event
    - Composition Benchmark: +25.8 bps/event
    - Selection Value: -10.6 bps/event (Agent does not outperform random)];
    G --> H[Recommendation: Report Deployment Value & Event-Selection Value Separately];