Problem
Our consumer products needed a personalized feed that stayed fast under real production traffic.
Hard decisions
- Fan-out, asyncFeeds are built by an asynchronous fan-out pipeline instead of being assembled on each request.
- Cache processed feeds in RedisServing pre-processed feeds from Redis is what holds response times under 50 ms in production.
- Fix the read path, not just the symptomAfter the outage: paginated queries and database-level indexes instead of unbounded collection reads, then a sweep of the codebase for the same pattern.
What broke
The feed service started timing out, but only for our largest accounts. I traced it to unbounded collection queries and N+1 lookups, redesigned the read path with paginated queries and database-level indexes, and brought tail latency back under SLA the same day.
Then I swept the codebase for the same anti-pattern and found two more services with it before they turned into incidents.