Machine Learning System Design Interview #4 - The Infinite Stream Trap
Why batch thinking fails in infinite streams - and how Reservoir Sampling saves you
Youโre in a Senior ML System Design interview at Twitter. The interviewer sets a trap:
โWe have a firehose of tweets coming in at 50k TPS. I need you to maintain a statistically representative sample of exactly 10,000 tweets for a training buffer at all times. The stream never stops. You cannot store the full history.โ
90% of candidates walk right into the wall.
Most candidates revert to ๐๐๐ญ๐๐ก ๐๐ก๐ข๐ง๐ค๐ข๐ง๐ .
They say: โEasy. Iโll buffer the last hour of data into S3, load it into a Dataframe, and run ๐ฅ๐ง.๐ด๐ข๐ฎ๐ฑ๐ญ๐ฆ(๐ฏ=10000)โ . Or they suggest:
โI will just flip a coin and keep every 100th tweet.โ
The interviewer stops you. โYou just crashed production.โ
The Reality:


