Paper: arXiv 2610.10407
Authors: Yizhen Xie, Mengyang Liu
Abstract
As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.
Complexity vs Empirical Score
- Math Complexity: 6.5/10
- Empirical Rigor: 8.0/10
- Quadrant: Holy Grail — high math complexity, high empirical rigor
Why this score: This paper presents a novel approach to option trading using LLM agents, addressing the complex decision space effectively. The empirical results are strong and well-documented, demonstrating significant outperformance against baselines. The methodology is clearly outlined, suggesting good reproducibility.
Research Flowchart
flowchart TD
A[Research Goal: Develop Agentic System for Optimal Option Trading] --> B{Key Methodology Steps};
B --> C[Data/Inputs: Historical Option Data, News, LLM (Qwen3.8-27B)];
C --> D[Computational Processes: SFT, RL (on Qwen3.8-27B)];
D --> E[SOTA: Structured Option-Strategy Selection Framework];
E --> F{Key Findings/Outcomes};
F --> G[18.3% Total Return, Sharpe Ratio 1.60, Max Drawdown 8.96%];
F --> H[News improves teacher trajectories, but hurts RL OOS return];