MERN · Python · Celery · AWS EC2 · AI
PodAgain: AI Learning & Podcast Platform
Podcasts, blogs, PDFs and documents turned into a searchable knowledge base, then into a two-person AI podcast. Rebuilt from a blocking prototype into a queue-based platform with GPU servers that start and stop on demand.
Screens are illustrative layouts of the product.
Overview
What it is
PodAgain is an AI-powered learning platform that turns long-form content into interactive, conversational audio. Users add podcast RSS feeds, blogs, PDFs and documents. The platform processes and indexes them, so users can later ask questions or give a prompt and generate a two-person AI podcast based on what they have learned.
The problem
Why it was needed
The project arrived as a largely vibe-coded prototype. APIs were synchronous and blocking, so every expensive AI operation was tied to a single user request, and there was no real infrastructure for long-running AI workloads. The GPU server did far more than GPU work and ran around the clock.
Features
What We Built
Queue-based job system
Celery workers replaced blocking requests for ingestion, transcription, content processing, embeddings, vector indexing and audio generation.
Leaner GPU server
Business logic moved out of the Python AI server and into the main backend, so the expensive GPU instance only does GPU work.
Smart GPU scheduling
The AWS EC2 GPU instance starts when a job needs it and shuts down after a configurable idle period.
Open-source voices
ElevenLabs was replaced with Qwen-based generation for the conversational two-host podcast audio.
Credits and billing
Subscription tiers plus usage-based minute credits for transcription, AI voice and audio processing.
Retrieval pipeline
AssemblyAI transcripts and document text are embedded into Pinecone for semantic search across everything a user has added.
Frontend overhaul
A refreshed interface and a product that now goes well beyond its original podcast-only workflow.
Architecture
How It Fits Together
- FrontendReact app
- Node / Express backendauth, billing, business logic
- Job queueCelery + Redis
- Worker processesingest, transcribe, embed
- AI / GPU infrastructurestarted on demand
- Vector databasePinecone
Long-running workloads run independently of user-facing requests, and GPU spend follows actual usage.
Roadmap
From Idea to Launch
How the PodAgain work was sequenced, from audit to launch.
- 1
Audit the prototype
1 week- Map every blocking endpoint and where GPU time was spent
- List content sources and the target user flows
- Agree what moves out of the AI server
- 2
Queue architecture
1-2 weeks- Introduce Celery and Redis alongside the Node backend
- Define job types, retries and status reporting
- Move business logic from the Python AI server into the backend
- 3
Ingestion and retrieval
2-3 weeks- RSS, blog, PDF and document ingestion jobs
- AssemblyAI transcription and text processing
- Embedding generation and Pinecone indexing, with semantic search
- 4
GPU scheduling and audio
2 weeks- Start/stop automation for the EC2 GPU instance with an idle timer
- Swap ElevenLabs for Qwen two-host audio generation
- Prompt-to-podcast generation end to end
- 5
Credits, UI and launch
2 weeks- Subscription tiers and minute-based credits
- Frontend overhaul across the new content types
- Monitoring, load testing and production rollout
Want something like this?
Tell us what it should do and we'll come back with a stack, a scope and a timeline.