AI Infrastructure for Trading Systems

lesson 3 of 5

Low-latency inference and model serving

Inference is the moment a model is actually used: live data goes in, a decision or score comes out. In trading, the value of that output decays with time, so serving infrastructure is designed around a latency budget.

Latency budgets

A latency budget is the total time a system is allowed to take from an event to an action, broken down across every step: data arrival, feature computation, model inference, risk checks and order submission. Each component gets a share.

Averages are misleading here. Engineers track percentiles: the median (p50) describes a typical request, while the 99th percentile (p99) describes the slow tail. In markets, the slow tail tends to arrive exactly when activity spikes, which is when speed matters most.

Making models fast enough

  • Right-sizing: the smallest model that does the job well is usually the best production choice.
  • Quantisation: storing model weights at lower numerical precision to reduce memory and speed up inference.
  • Distillation: training a smaller model to reproduce the behaviour of a larger one.
  • Hardware choice: GPUs suit large, parallel workloads; well-optimised CPU inference is often enough for smaller models.
  • Caching and precomputation: computing what can be known in advance before the time-critical moment arrives.

Scale and context

Institutional high-frequency trading measures latency in microseconds and places servers physically next to exchange matching engines, which is known as co-location. Most retail and systematic strategies act on bar closes measured in seconds or minutes. The principles are the same; the budget is different. Knowing which world a strategy lives in prevents both over-engineering and under-engineering.

Educational content only. Not financial advice.