Completed
Strata

Strata

A production-grade distributed social platform featuring realtime notifications, activity feeds, Redis caching, BullMQ workers, Dead Letter Queues, monitoring dashboards, and horizontally scalable Socket.IO infrastructure.

Overview

Strata is a backend-first social platform built to answer one question: what breaks when notifications, feeds, and presence all go realtime? Constraint was a single Postgres + Redis pair serving 3 API instances. Notifications fan out through BullMQ workers with a Dead Letter Queue and replay dashboard; Socket.IO rooms sync over the Redis adapter with sticky sessions; counts and feed pages sit behind Redis caching with explicit invalidation. Tested to 500 concurrent sockets with notification p95 under 400ms.

What Users Can Do

  • Create accounts with JWT sessions, post, like, comment, and follow/unfollow.
  • View a paginated feed built from the follow graph.
  • Receive realtime notifications for follows, likes, and comments with instant unread counts.
  • See online presence and last-seen status in realtime.
  • Inspect queues, workers, health, and replay history in Bull Board.

Why I built this

  • To move past CRUD and practice reliability engineering: retries, DLQs, replay, and recovery metrics.
  • Tradeoff: fan-out-on-write for notifications (fast reads, bursty writes) vs fan-out-on-read for feeds (cheap writes, slower reads) - documented per path.
  • Tradeoff: Redis-cached notification counts over live COUNT(*) queries - 60% fewer DB reads at the cost of explicit invalidation on every write.
  • Tradeoff: Socket.IO + Redis adapter with sticky sessions over bare WebSockets - room sync across instances at the cost of LB affinity.
  • To build ops muscle: queue metrics, worker metrics, and replay tooling, not just features.

Tech Stack

Next.js
TypeScript
PostgreSQL
Prisma
Redis
BullMQBullMQ
Socket.IO
Turborepo

After launch & Impact

  • Event-driven notification pipeline (BullMQ + Redis Pub/Sub) delivers p95 under 400ms at 500 concurrent sockets across 3 API instances.
  • Dead Letter Queue with replay dashboard recovered 98%+ of poison notification jobs in 1k-job chaos tests (5 attempts, exponential backoff, then DLQ).
  • Redis caching for notification counts and feed pages hit ~87% and cut repeated DB queries by ~60%.
  • Presence tracking via Redis TTLs survived rolling restarts with zero ghost-online users after 60s reconciliation.
  • Redis-blip test (30s outage): zero lost notifications - jobs retried with backoff and drained with queue-lag p95 under 12s on recovery.
  • Socket disconnect test: clients resync missed notifications on reconnect; honest limit is single-region Redis as SPOF, documented next step is Sentinel/cluster.

Future Plans

  • Ship fan-out-on-write feed precomputation with timeline cache warming.
  • Put 3 API instances behind a load balancer with sticky-session demo and failover notes.
  • Add media upload + moderation pipeline as a second BullMQ flow.
  • Build a full ops dashboard for notifications, feeds, and recovery rate.