Paper: arXiv 2609.32848
Authors: Vincent Maciejewski
Abstract
HFT systems are conventionally built as a single-threaded event loop, on the rule that every thread hop adds latency. We test that rule against a measurement study of more than a year of CME market data for the NQ front-month contract, following every packet and matching-engine transaction through the feed’s two exchange timestamps, and checking the results against a live production receiver. Packets arrive in near-critical self-exciting clusters that belong to the matching engine’s transactions, not to how the exchange packs them. The engine often processes consecutive transactions within a fraction of a microsecond, while the market-data publisher sends at most one packet per publisher period of about 7.5 microseconds, so a burst reaches the receiver as a train of packets one period apart. This yields design principles for HFT systems. First, a receiver that handles each packet within one publisher period never queues on arrivals, however bursty the market; there one thread is best. Second, above that period a queueing tail appears, driven by the timing of transactions, not by packet rate or size, and two threads can be better than one: splitting the servicing chain into two stages on separate threads removes most of the tail at the cost of one hop on the median. Third, only the slowest stage matters, so a split pays only if it shortens it. Fourth, just under the period, where the production receiver runs, the remaining tail comes from multi-message packets and variable service times, and the levers are cost per message and spread of service, not thread count. An analytic framework, a burst-limit throughput identity and an exact reduction of the tandem to a single bottleneck server, supports these results.
Complexity vs Empirical Score
- Math Complexity: 7.5/10
- Empirical Rigor: 9.0/10
- Quadrant: Holy Grail — high math complexity, high empirical rigor
Why this score: This paper presents a highly rigorous empirical study of HFT system design principles, backed by a sophisticated analytical framework. Its novelty lies in challenging conventional wisdom regarding single-threaded event loops through detailed measurement and simulation. The comprehensive analysis and clear recommendations make it a significant contribution.
Research Flowchart
flowchart TD
A[Research Goal: Validate HFT Event Loop Rule vs. Latency] --> B{Key Methodology: Measurement Study of CME Market Data};
B --> C[Data/Inputs: >1 Year NQ Front-Month CME Market Data, Live Production Receiver Logs];
C --> D{Computational Process: Packet/Transaction Tracking, Latency Analysis, Queueing Models};
D --> E[Key Finding 1: Packets arrive in near-critical self-exciting clusters belonging to matching engine transactions];
E --> F[Key Finding 2: HFT Design Principles based on Publisher Period & Transaction Timing];
F --> G[Key Outcome: Analytic Framework, Burst-Limit Throughput Identity, Tandem Reduction for HFT System Design];