lesson 2 of 5
Data pipelines and feature stores
Every model is only as good as the data that reaches it. In trading, data problems are rarely dramatic. They are small, silent and expensive: a missing bar, a timestamp in the wrong zone, a contract roll that shifts every price on a continuous chart.
Time is the hardest part
- Exchange time versus receive time: when an event happened at the exchange is not the same as when your system learned about it.
- Bar boundaries: a one-minute bar is only final when the minute closes. Using it earlier is lookahead.
- Time zones and daylight saving: sessions defined in local market time move against UTC twice a year.
- Exchange holiday calendars: shortened sessions change what 'normal' volume and volatility look like.
Point-in-time correctness
A point-in-time dataset answers one question honestly: what did the system know at this exact moment? Every feature used in training must be computed only from data available at the time it would have been used live. Breaking this rule is the infrastructure version of lookahead bias, and it produces models that look excellent in testing and disappoint in production.
Feature stores
A feature store is a shared system for defining, computing and serving features. Its main job is to make sure a feature is calculated the same way in training and in live inference. When the two drift apart, the result is called training-serving skew: the model was trained on one version of reality and is being asked to act on another.
Well-run teams define each feature once, version it, and serve it to both the training pipeline and the live system from the same definition.
Educational content only. Not financial advice.