Volume-weighted average price is the sum of price times volume over a window divided by total volume: the average price at which the market actually traded. It is the dominant execution benchmark because it is cheap to compute, hard for a broker to dispute, and an order that fills at the day’s VWAP has by definition done as well as the average participant. The catch is that VWAP is only known at the close, so tracking it means predicting the market’s intraday volume profile and spreading your order in proportion. A VWAP model is therefore a volume model plus a scheduler, and the two are evaluated differently.
Scope: one corpus paper carries the exact tag “VWAP prediction”. The ranked list below widens to VWAP-execution and volume-forecasting tags, where the methods live; sixteen papers qualify.
The volume-profile problem
Intraday equity volume follows a U-shape: heavy at the open, thin at midday, heavy into the close. The simplest VWAP schedule is a static profile: average each stock’s historical share of daily volume per time bucket over a rolling window of 20 to 60 days and trade that fraction of the order in each bucket. It is a hard baseline because the shape is stable; what it misses is day-specific deviation (news, rebalances, expiries) and the scaling of the curve by that day’s total volume.
Dynamic profiles update the remaining schedule as realised volume arrives. The decomposition that most of the better papers use is daily total times intraday ratio; the ratio is forecastable from the historical shape, the total is forecastable from recent days and calendar features, and errors in the two compound. Forecasting Intraday Volume in Equity Markets with Machine Learning (rigor 7.5) finds intraday volume highly predictable once a machine-learning suite with high-frequency predictors exploits commonality, the cross-sectional co-movement of volume, and quantifies the gain through VWAP-strategy economics rather than forecast error alone. IVE (rigor 8.5) predicts the one-minute volume ratio with an encoder-decoder Transformer and a distribution head, so the scheduler gets a mean and a standard deviation and can anticipate volume spikes; it is one of the few papers here with a live test, two and a half months in Korea against VWAP benchmarks.
Two papers treat the profile as a statistical object rather than a regression target. Modelling financial volume curves with hierarchical Poisson processes (math 8.5) models trade arrival times as inhomogeneous Poisson processes with a hierarchical Dirichlet-process prior, so stocks share shape information within industries, demonstrated on WRDS TAQ data including Apple. Stock Volume Forecasting with Advanced Information by Conditional Variational Auto-Encoder works at the daily horizon, conditioning generated volume paths on known future events such as rebalancing dates and showing better correlation fit than linear baselines on EURO STOXX 50 names. Both are closer to scenario generation than to a production forecaster.
The paper tagged exactly “VWAP prediction”, Enhancing OHLC Data with Timing Features (rigor 8.0), is a different task again: predicting a bar’s VWAP from OHLC bars enriched with the timestamps of the open, high, low and close, on Russell 3000 names in 2021 with temporal splits. A feature-engineering result about bar data, not an execution scheduler.
From profile to schedule
Once you have a profile, the classical scheduler trades it. The newer literature argues the forecast is the wrong intermediate target. Deep Learning for VWAP Execution in Crypto Markets: Beyond the Volume Curve optimises allocation directly against VWAP slippage and reports that the allocations which minimise slippage diverge from the allocations an accurate volume forecast would imply. An important idea with a weak evidence base: rigor 4.0 for thin reporting. The same author’s follow-ups are stronger. Recurrent Neural Networks for Dynamic VWAP Execution (rigor 8.0) adds recurrent state and a dynamic adjustment mechanism and reports 10–15% execution-performance gains in liquid crypto markets; VWAP Execution with Signature-Enhanced Transformers (rigor 8.0) trains one model across 80 hourly crypto pairs and shows the gains persist on assets outside the training set; LEMs generalises to flexible horizons and fixed-notional versus fixed-quantity orders, tested on crypto and Dow constituents. Read the chain as one research programme on hourly crypto data; whether it transfers to equity minute bars is untested in the corpus.
The reinforcement-learning branch keeps the profile and learns the tactics. Hierarchical Deep RL for VWAP Strategy Optimization (rigor 7.5) allocates tranches from an LSTM volume forecast, then lets lower-level agents place within each tranche, reporting an average 1.16 bps saving over the best baseline on Shanghai-listed stocks. An Adaptive Dual-level RL Approach for Optimal Trade Execution pairs a Transformer for the global U-shape with an LSTM for local placement under PPO, without published data or code. RL-Exec (rigor 8.5) is the evaluation model to copy: a PPO liquidator on BTC-USD book replays with transient impact, fees and latency, compared against TWAP and a book-liquidity VWAP on identical timestamps, with Wilcoxon tests and FDR correction, showing a gap that grows from 2–3 bps at 30 minutes to 23 bps at 120 minutes. Its test set is one month. Diverse Approaches to Optimal Execution Schedule Generation (rigor 8.0) reports 2.13 bps arrival slippage for a CNN-PPO scheduler against 5.23 bps for VWAP across 4,900 out-of-sample US equity orders, inside a calibrated transient-impact simulator. Model-Free Passive Execution via Order-Level Shadowing is the other extreme: no forecast, placement copied from observed order flow, offered as a model-free benchmark on a year of CME ES replays with open-source code.
How to evaluate a VWAP model
Point accuracy of the volume forecast is not the objective, and the “Beyond the Volume Curve” result is the reason: the best forecast and the best schedule differ. Judge a VWAP paper on these instead:
- Tracking-error distribution, not its mean. Slippage against realised VWAP in basis points, per order, with its standard deviation, tails and sign. Zero mean slippage with a fat left tail is worse for a desk than a small positive mean with tight dispersion. Absolute and quadratic VWAP loss, as used in the signature-Transformer paper, are the right family; MAPE on the volume ratio is not.
- Paired comparison on identical conditions. Same orders, same timestamps, same fee and impact model for every strategy. RL-Exec’s per-day protocol with ten start times aggregated to a daily score is a template for avoiding pseudo-replication.
- Horizon and liquidity dependence. Gains at two hours mean little for a 15-minute order, and gains in liquid names may vanish in the tail of the universe; ask for results sliced both ways.
- Self-impact. Your own fills are part of the VWAP you are measured against. A model tested only at participation rates where this is negligible has not been tested at the sizes VWAP algorithms are used for. The transaction-cost guide has the impact arithmetic.
- Out-of-sample assets and periods. The better papers have both (signature Transformers across assets, RL-Exec across time); single-asset, single-year results are common elsewhere.
- The baseline. A static 20-day profile and a POV scheduler are the right dumb baselines; comparing only against TWAP flatters any volume-aware method.
Everything in the RL-evaluation checklist and the microstructure-paper guide applies to the learned schedulers above.
- IVE: Enhanced Probabilistic Forecasting of Intraday Volume Ratio with Transformers Holy Grail Rigor 8.5 Math 7
- VWAP Execution with Signature-Enhanced Transformers: A Multi-Asset Learning Approach Holy Grail Rigor 8 Math 7.5
- Recurrent Neural Networks for Dynamic VWAP Execution: Adaptive Trading Strategies with Temporal Kolmogorov-Arnold Networks Holy Grail Rigor 8 Math 7.5
- Optimal Execution in Intraday Energy Markets under Hawkes Processes with Transient Impact Holy Grail Rigor 7 Math 8.5
- Stock Volume Forecasting with Advanced Information by Conditional Variational Auto-Encoder Holy Grail Rigor 7 Math 8.5
- Modelling financial volume curves with hierarchical Poisson processes Holy Grail Rigor 7 Math 8.5
- An Adaptive Dual-level Reinforcement Learning Approach for Optimal Trade Execution Holy Grail Rigor 7 Math 8.5
- RL-Exec: Impact-Aware Reinforcement Learning for Opportunistic Optimal Liquidation, Outperforms TWAP and a Book-Liquidity VWAP on BTC-USD Replays Holy Grail Rigor 8.5 Math 6
- LEMs: A Primer On Large Execution Models Holy Grail Rigor 7 Math 7.5
- Forecasting Intraday Volume in Equity Markets with Machine Learning Holy Grail Rigor 7.5 Math 6
- Hierarchical Deep Reinforcement Learning for VWAP Strategy Optimization Holy Grail Rigor 7.5 Math 6
- An Optimal Control Strategy for Execution of Large Stock Orders Using LSTMs Holy Grail Rigor 6.8 Math 6.5
- Enhancing OHLC Data with Timing Features: A Machine Learning Evaluation Street Traders Rigor 8 Math 4.5
- Model-Free Passive Execution via Order-Level Shadowing Street Traders Rigor 8 Math 4
- Deep Reinforcement Learning for Optimum Order Execution: Mitigating Risk and Maximizing Returns Holy Grail Rigor 5 Math 6.5
- Deep Learning for VWAP Execution in Crypto Markets: Beyond the Volume Curve Lab Rats Rigor 4 Math 6
Where this leaves a practitioner
For equities, the strongest corpus evidence is that intraday volume is predictable with commonality features and that probabilistic one-minute ratio forecasts can beat VWAP in live trading. For crypto, the direct-optimisation and multi-asset results are promising but come from one programme on hourly data. On evaluation, the RL execution papers lead: paired, cost-matched slippage distributions with inference, which is what a desk needs before replacing a scheduler.
Related reading: Transaction costs, slippage and impact → · Evaluate a microstructure paper → · Evaluate an RL trading paper → · Market-making mechanics → · Hubs: HFT & Optimal Execution · Market Microstructure · Reinforcement Learning
Broader area: Machine Learning · All topics: research topics → · Full archive: every paper →