Completed

TranscodeX
A distributed video processing and streaming platform featuring multipart uploads, FFmpeg-powered HLS transcoding, AI-generated captions, transcript-based chat, realtime processing updates, and adaptive bitrate streaming.
Overview
Transcodex is a distributed video pipeline (upload → transcode → stream) built under single-region, single-box constraints: a 4 vCPU worker host, S3-compatible storage, and Redis as the only shared state. The API never touches FFmpeg - it takes multipart uploads and enqueues jobs; BullMQ workers own transcoding, thumbnails, captions, and progress fanout via Socket.IO + Redis Pub/Sub. Tested to 4 concurrent 1080p encodes with API p95 under 120ms while workers were pegged.
What Users Can Do
- Upload multi-GB video files using S3 multipart uploads with resume on failure.
- Track transcode progress in realtime over Socket.IO without polling.
- Watch adaptive HLS with manual 480p / 720p / 1080p switching, speed control, and subtitles.
- Get Whisper-generated captions and searchable transcripts per video.
- Chat with a transcript-grounded AI assistant about video content.
- Browse auto-generated thumbnails for every processed video.
Why I built this
- To learn how YouTube-style systems separate upload, processing, and delivery - and where they break under load.
- Tradeoff: S3 multipart + server-coordinated parts over direct-to-S3 POSTs, chosen for resume and progress accuracy at the cost of API coordination.
- Tradeoff: BullMQ + Redis queues over in-process workers, chosen for retries, DLQ, and horizontal worker scale-out.
- Tradeoff: Socket.IO + Redis Pub/Sub over polling/SSE, chosen for room-based progress fanout across API instances.
- To push AI past chatbots: captioning, transcript extraction, and grounded Q&A in one pipeline.
Tech Stack
Next.js
TypeScript
PostgreSQL
Prisma
Redis
BullMQ
After launch & Impact
- Multipart uploads in 8 MB parts lifted large-file completion from ~70% (single PUT) to 99%+ across 50+ test uploads up to 5 GB.
- BullMQ pipeline sustains 4 concurrent 1080p encodes on a 4 vCPU box - ~3.5 min per 500 MB source to 480p/720p/1080p HLS, queue-lag p95 under 8s at 20 queued jobs.
- FFmpeg workers capped at 2 GB RSS with concurrency limits: OOM kills dropped to zero across 100 test encodes; poison jobs retry 3x then land in DLQ with 100% replay success via Bull Board.
- Realtime progress fanout every 2s survived API restarts (Redis-backed state) with zero stuck progress bars in disconnect tests.
- Whisper captions + Gemini transcript chat answered grounded questions with citations; API held p95 under 120ms at 200 rps while workers were saturated.
- Single-region transcode cost ~$0.11/hr on Hetzner 4 vCPU vs ~$0.45/hr managed - documented limit: no GPU, no CDN, cold starts on first segment.
Future Plans
- Add CDN + signed HLS URLs for global delivery.
- Introduce GPU-accelerated transcoding and per-title bitrate ladders.
- Add vector search over transcripts with timestamp-linked AI answers.
- Deploy autoscaled worker pools with queue-lag-based scaling.