This hub holds the mathematical engine room of quantitative finance: stochastic control (choose a policy that steers a diffusion), optimal stopping (choose a time to act), and the machinery that solves them — Hamilton–Jacobi–Bellman equations, viscosity solutions, backward stochastic differential equations, and mean-field games for the many-player limit. The classic applications are Merton-style consumption and investment, optimal liquidation, dividend and reinsurance control, and American-option exercise.
Most papers here are Lab Rats by our scoring: deep theory, little or no data. That is not a criticism of the work, but it changes how you read it. The questions that matter are whether the model’s state variables are observable in practice, whether the closed-form or numerical solution degrades gracefully when parameters are estimated rather than known, and whether a paper that claims a strategy ever confronts it with transaction costs or discrete rebalancing. The handful of papers that pair a control result with a calibrated numerical study rank highest below.
Related hubs: Options & Derivatives, Portfolio Optimization, HFT & Optimal Execution, Insurance & Actuarial Risk.
We develop policy gradients methods for stochastic control with exit time in a model-free setting. We propose two types of algorithms for learning either directly the optimal policy or by learning alternately the value function (critic) and the optimal control (actor). The use of randomized policies
Stochastic Gradient Descent Langevin Dynamics (SGLD) algorithms, which add noise to the classic gradient descent, are known to improve the training of neural networks in some cases where the neural network is very deep. In this paper we study the possibilities of training acceleration for the numeri
We present an approximate dynamic programming framework for designing degradation-aware market participation policies for battery energy storage systems. The approach employs a tailored value function approximation that reduces the state space to state of charge and battery health, while performing
This paper investigates a time-inconsistent portfolio selection problem in the incomplete mar ket model, integrating expected utility maximization with risk control. The objective functional balances the expected utility and variance on log returns, giving rise to time inconsistency and motivating t
We study optimal consumption and retirement using a Cobb-Douglas utility and a simple model in which an interesting bifurcation arises. With high wealth, individuals plan to retire. With low wealth they plan to never retire. At a critical level of initial wealth they may choose to defer this decisio
We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility function and price process. We first consider a relaxed versio
High-dimensional option pricing and hedging present significant challenges in quantitative finance, where traditional PDE-based methods struggle with the curse of dimensionality. The BSDE framework offers a computationally efficient alternative to PDE-based methods, and recently proposed deep BSDE s
We propose and analyze a continuous-time robust reinforcement learning framework for optimal stopping under ambiguity. In this framework, an agent chooses a robust exploratory stopping time motivated by two objectives: robust decision-making under ambiguity and learning about the unknown environment
Traditional stochastic control methods in finance rely on simplifying assumptions that often fail in real world markets. While these methods work well in specific, well defined scenarios, they underperform when market conditions change. We introduce FinFlowRL, a novel framework for financial stochas
We propose two signature-based methods to solve the optimal stopping problem - that is, to price American options - in non-Markovian frameworks. Both methods rely on a global approximation result for $L^p-$functionals on rough path-spaces, using linear functionals of robust, rough path signatures. I
We propose a novel computational procedure for quadratic hedging in high-dimensional incomplete markets, covering mean-variance hedging and local risk minimization. Starting from the observation that both quadratic approaches can be treated from the point of view of backward stochastic differential
This paper studies a joint stochastic optimal control and stopping (JCtrlOS) problem motivated by aquaculture operations, where the objective is to maximize farm profit through an optimal feeding strategy and harvesting time under stochastic price dynamics. We introduce a simplified aquaculture mode
We consider a Bayesian diffusion control problem of expected terminal utility maximization. The controller imposes a prior distribution on the unknown drift of an underlying diffusion. The Bayesian optimal control, tracking the posterior distribution of the unknown drift, can be characterized explic
An agent-based modelling methodology for the joint price evolution of two stocks is put forward. The method models future multidimensional price trajectories reflecting how a class of agents rebalance their portfolios in an operational way by reacting to how stocks’ charts unfold. Prices are express
As the developed world replaces Defined Benefit (DB) pension plans with Defined Contribution (DC) plans, there is a need to develop decumulation strategies for DC plan holders. Optimal decumulation can be viewed as a problem in optimal stochastic control. Formulation as a control problem requires sp
This paper studies the ubiquitous problem of liquidating large quantities of highly correlated stocks, a task frequently encountered by institutional investors and proprietary trading firms. Traditional methods in this setting suffer from the curse of dimensionality, making them impractical for high
Previous research on option strategies has primarily focused on their behavior near expiration, with limited attention to the transient value process of the portfolio. In this paper, we formulate Iron Condor portfolio optimization as a stochastic optimal control problem, examining the impact of the
A corporate bond trader in a typical sell side institution such as a bank provides liquidity to the market participants by buying/selling securities and maintaining an inventory. Upon receiving a request for a buy/sell price quote (RFQ), the trader provides a quote by adding a spread over a \textit{
In this paper, we propose a multidimensional statistical model of intraday electricity prices at the scale of the trading session, which allows all products to be simulated simultaneously. This model, based on Poisson measures and inspired by the Common Shock Poisson Model, reproduces the Samuelson
We propose an approach to applying neural networks on linear parabolic variational inequalities. We use loss functions that directly incorporate the variational inequality on the whole domain to bypass the need to determine the stopping region in advance and prove the existence of neural networks wh