Market data streaming platform
A Kafka pipeline aggregating market ticks into one-second bars, with a live dashboard, batched Snowflake storage, and latency monitoring.
Python · Kafka · Snowflake · Streamlit
The problem
Live monitoring and historical analysis need the same data, but they have different latency and cost requirements. Sending every tick to a warehouse is an expensive answer to the wrong question.
What I built
- Separate ingestion, aggregation, and persistence with Kafka topics and explicit data contracts.
- Aggregate market ticks into one-second OHLCV bars for the local dashboard, while batching rollups for Snowflake.
- Expose ingestion lag and data freshness through timestamped logging, metrics, and dashboard views.
- Use validation scripts and deterministic identifiers to reason about duplicates and aggregate correctness.
Key tradeoff
Freshness versus compute cost. The real-time path serves immediate monitoring; batched warehouse writes support analysis without making every incoming event a warehouse operation.