Paper: arXiv 2609.33767

Authors: Wee Ling Tan, Stephen Roberts, Stefan Zohren

Abstract

We present an end-to-end deep learning framework for systematic options trading that directly embeds hedging behavior through explicit control of portfolio-level risk exposures. While neural networks trained to optimize risk-adjusted performance have been shown to outperform traditional rules-based strategies, such approaches remain agnostic to the sensitivities of the resulting portfolios with respect to specific underlying risk factors. We propose a general training objective that combines a performance-driven loss with a differentiable risk-sensitivity penalty, enforcing neutrality to selected risk dimensions. Unlike reinforcement learning methods that approximate optimal hedging policies via simulated market dynamics, our framework operates entirely on historical data and jointly optimizes risk-adjusted returns and targeted risk constraints in a single learning problem. We instantiate the framework on static delta-neutral straddle portfolios with the penalty directed at first-order directional exposure, and evaluate two penalty variants – an exposure-normalized penalty and a Greek-ratio drift penalty. Empirical results on Nasdaq 100 equity options demonstrate that appropriately calibrated regularization simultaneously improves out-of-sample risk-adjusted performance relative to an unregularized baseline while reducing realized directional exposure.

Complexity vs Empirical Score

  • Math Complexity: 6.5/10
  • Empirical Rigor: 8.0/10
  • Quadrant: Holy Grail — high math complexity, high empirical rigor

Why this score: This paper presents a novel deep learning framework for options trading that integrates risk-sensitivity penalties directly into the learning objective, moving beyond traditional risk-agnostic approaches. The empirical evaluation on Nasdaq 100 equity options demonstrates strong out-of-sample performance and effective risk control, indicating a robust methodology. The combination of advanced mathematical concepts in deep learning and rigorous empirical validation places it firmly in the ‘Holy Grail’ quadrant.

Research Flowchart

  flowchart TD
    A[Research Goal: Tame Greeks in Option Portfolios] --> B{Key Problem: Traditional NN agnostic to risk sensitivities};
    B --> C[Methodology: Deep Learning Framework with Differentiable Risk-Sensitivity Penalty];
    C --> D[Inputs: Historical Options Data (Nasdaq 100 Equities)];
    D --> E[Computational Process: Joint Optimization of Performance & Targeted Risk (Delta-Neutral Straddle Portfolios)];
    E --> F{Outcomes: Improved Risk-Adjusted Performance & Reduced Realized Directional Exposure};