AI Interview Prep

AI Interview Prep

Machine Learning System Design Interview #19 - The Database-as-Queue Trap

How relying on a persistence layer for real-time inference creates silent bottlenecks, and how event brokers fix them.

Hao Hoang's avatar
Hao Hoang
Dec 05, 2025
โˆ™ Paid

Youโ€™re in a Machine Learning System Design interview at Google. The interviewer sets a trap:

โ€œWe have 3 upstream microservices generating features. They write to a central ๐˜—๐˜ฐ๐˜ด๐˜ต๐˜จ๐˜ณ๐˜ฆ๐˜ด ๐˜‹๐˜‰. Your ML Service queries that DB to get the input vector for inference. How do we scale this to 50k requests per second?โ€

90% of candidates walk right into the trap. They start optimizing the SQL.

They immediately focuses on the database performance. They suggest:

- Adding aggressive Read Replicas to handle the load.

- Implementing complex caching layers (Redis) to offload the DB.

- Sharding the Postgres instance based on UserID.

It feels right. ๐˜๐˜ง ๐˜ต๐˜ฉ๐˜ฆ ๐˜ฅ๐˜ข๐˜ต๐˜ข๐˜ฃ๐˜ข๐˜ด๐˜ฆ ๐˜ช๐˜ด ๐˜ด๐˜ญ๐˜ฐ๐˜ธ, ๐˜ง๐˜ช๐˜น ๐˜ต๐˜ฉ๐˜ฆ ๐˜ฅ๐˜ข๐˜ต๐˜ข๐˜ฃ๐˜ข๐˜ด๐˜ฆ, ๐˜ณ๐˜ช๐˜จ๐˜ฉ๐˜ต?

But they arenโ€™t solving a ๐๐ฎ๐ž๐ซ๐ฒ ๐Ž๐ฉ๐ญ๐ข๐ฆ๐ข๐ณ๐š๐ญ๐ข๐จ๐ง problem. They are solving a ๐‚๐จ๐ฎ๐ฉ๐ฅ๐ข๐ง๐  problem.

AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
ยฉ 2026 Hao Hoang ยท Privacy โˆ™ Terms โˆ™ Collection notice
Start your SubstackGet the app
Substack is the home for great culture