{"slug": "how-i-built-an-ai-studio-that-turns-manhwa-chapters-into-narrated-recap-videos", "title": "How I Built an AI Studio That Turns Manhwa Chapters into Narrated Recap Videos", "summary": "A developer built AniFlow, an open-source (MIT) self-hosted AI studio that converts manhwa and webtoon chapters into finished narrated recap videos. The pipeline chains YOLOv8 panel detection, speech-bubble detection and OCR, inpainting-based bubble removal, local LLM narration via Ollama, multi-voice TTS, and FFmpeg compositing, running fully locally with no API keys or subscriptions. The developer said the hardest problems were cross-art-style panel detection and artifact-free bubble inpainting, and advised investing in training-data diversity and an evaluation harness earlier.", "body_md": "If you've ever fallen down the rabbit hole of manhwa recap channels on YouTube — \"The Weakest Hunter Becomes the Strongest...\" — you know the format: dramatic narration over panning comic panels, 10 minutes per video, millions of views. I kept wondering: could that entire pipeline be automated, running locally, with no subscriptions?\n\nThat's how AniFlow was born — a self-hosted AI studio that takes manhwa/webtoon chapters and renders finished recap videos. It's open source (MIT): github.com/aashish254/Aniflow\n\nHere's how the pipeline works, and what was hard about each step.\n\nThe pipeline\n\n1. Chapter ingestion. Paste a chapter URL or point at a local folder. The downloader grabs every page image in reading order.\n2. Panel detection (YOLOv8). Manhwa pages are vertical strips with irregular panel layouts. A trained YOLOv8 model detects panel boundaries so each panel can be cropped and sequenced for the video. This was the single hardest part — art styles vary wildly between series, and a detector trained on one artist's work falls apart on another's. Getting reliable detection across styles took the most iteration.\n3. Speech bubble detection + text extraction. A second detection pass finds speech bubbles and text regions. OCR pulls the dialogue out, which becomes the script for dialogue-recap mode.\n4. Bubble removal. For narrated mode, bubbles get removed and the art inpainted so panels look clean — like a friend telling you the story over the artwork. Doing this without leaving visible artifacts on detailed backgrounds was the second-hardest problem.\n5. Narration writing (LLM). The extracted story content goes to a local LLM (Ollama by default, with optional Gemini/Claude) with a prompt tuned for that dramatic recap-channel voice. It writes the narration script panel by panel.\n6. Voiceover (TTS). Multi-voice TTS — Edge TTS or Kokoro locally, ElevenLabs optionally — speaks the script. Different voices for narration vs. dialogue.\n7. Video rendering. Everything gets composited with FFmpeg: Ken Burns pan/zoom over panels, background music, watermarks, subtitles. Out comes an upload-ready vertical video.\nTwo modes\nDialogue recap — keeps the original speech bubbles visible; extracted dialogue becomes the audio. Closest to reading the chapter.\nNarrated recap — bubbles removed and inpainted; the AI narrates the story like a recap channel.\nFully local by default\nThe whole stack runs on your own machine: Ollama for narration, Edge TTS for voice, no API keys, no subscriptions, nothing leaves your computer. Cloud models are optional upgrades, not requirements. There's also a batch mode that processes 50+ chapters unattended — a full series to video overnight.\nWhat I'd do differently\nIf I started over, I'd spend more time on the training data for panel detection up front instead of iterating on model tweaks — data diversity beat architecture fiddling every time. I'd also build the evaluation harness earlier: a small set of \"golden\" chapters across art styles to regression-test every change against.\nTry it\nRepo: github.com/aashish254/Aniflow (MIT). There's a web UI, a v0.1.0 release, and tutorial videos in the README.\nIf you're into self-hosting, computer vision, or the manhwa scene — I'd love your feedback, and contributors are very welcome.", "url": "https://wpnews.pro/news/how-i-built-an-ai-studio-that-turns-manhwa-chapters-into-narrated-recap-videos", "canonical_source": "https://dev.to/aashish124/how-i-built-an-ai-studio-that-turns-manhwa-chapters-into-narrated-recap-videos-5enl", "published_at": "2026-09-29 22:02:41+00:00", "updated_at": "2026-09-29 22:16:36.682061+00:00", "lang": "en", "topics": ["computer-vision", "generative-ai", "ai-tools", "large-language-models", "ai-products"], "entities": ["AniFlow", "YOLOv8", "Ollama", "Edge TTS", "Kokoro", "ElevenLabs", "FFmpeg", "Gemini"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-built-an-ai-studio-that-turns-manhwa-chapters-into-narrated-recap-videos", "markdown": "https://wpnews.pro/news/how-i-built-an-ai-studio-that-turns-manhwa-chapters-into-narrated-recap-videos.md", "text": "https://wpnews.pro/news/how-i-built-an-ai-studio-that-turns-manhwa-chapters-into-narrated-recap-videos.txt", "jsonld": "https://wpnews.pro/news/how-i-built-an-ai-studio-that-turns-manhwa-chapters-into-narrated-recap-videos.jsonld"}}