cd /news/computer-vision/how-i-built-an-ai-studio-that-turns-… Β· home β€Ί topics β€Ί computer-vision β€Ί article
[ARTICLE Β· art-142092] src=dev.to β†— pub= topic=computer-vision verified=true sentiment=↑ positive

How I Built an AI Studio That Turns Manhwa Chapters into Narrated Recap Videos

A developer built AniFlow, an open-source (MIT) self-hosted AI studio that converts manhwa and webtoon chapters into finished narrated recap videos. The pipeline chains YOLOv8 panel detection, speech-bubble detection and OCR, inpainting-based bubble removal, local LLM narration via Ollama, multi-voice TTS, and FFmpeg compositing, running fully locally with no API keys or subscriptions. The developer said the hardest problems were cross-art-style panel detection and artifact-free bubble inpainting, and advised investing in training-data diversity and an evaluation harness earlier.

by read3 min views6 publishedSep 29, 2026

If you've ever fallen down the rabbit hole of manhwa recap channels on YouTube β€” "The Weakest Hunter Becomes the Strongest..." β€” you know the format: dramatic narration over panning comic panels, 10 minutes per video, millions of views. I kept wondering: could that entire pipeline be automated, running locally, with no subscriptions? That's how AniFlow was born β€” a self-hosted AI studio that takes manhwa/webtoon chapters and renders finished recap videos. It's open source (MIT): github.com/aashish254/Aniflow

Here's how the pipeline works, and what was hard about each step.

The pipeline

  1. Chapter ingestion. Paste a chapter URL or point at a local folder. The down grabs every page image in reading order.
  2. Panel detection (YOLOv8). Manhwa pages are vertical strips with irregular panel layouts. A trained YOLOv8 model detects panel boundaries so each panel can be cropped and sequenced for the video. This was the single hardest part β€” art styles vary wildly between series, and a detector trained on one artist's work falls apart on another's. Getting reliable detection across styles took the most iteration.
  3. Speech bubble detection + text extraction. A second detection pass finds speech bubbles and text regions. OCR pulls the dialogue out, which becomes the script for dialogue-recap mode.
  4. Bubble removal. For narrated mode, bubbles get removed and the art inpainted so panels look clean β€” like a friend telling you the story over the artwork. Doing this without leaving visible artifacts on detailed backgrounds was the second-hardest problem.
  5. Narration writing (LLM). The extracted story content goes to a local LLM (Ollama by default, with optional Gemini/Claude) with a prompt tuned for that dramatic recap-channel voice. It writes the narration script panel by panel.
  6. Voiceover (TTS). Multi-voice TTS β€” Edge TTS or Kokoro locally, ElevenLabs optionally β€” speaks the script. Different voices for narration vs. dialogue.
  7. Video rendering. Everything gets composited with FFmpeg: Ken Burns pan/zoom over panels, background music, watermarks, subtitles. Out comes an upload-ready vertical video. Two modes Dialogue recap β€” keeps the original speech bubbles visible; extracted dialogue becomes the audio. Closest to reading the chapter. Narrated recap β€” bubbles removed and inpainted; the AI narrates the story like a recap channel. Fully local by default The whole stack runs on your own machine: Ollama for narration, Edge TTS for voice, no API keys, no subscriptions, nothing leaves your computer. Cloud models are optional upgrades, not requirements. There's also a batch mode that processes 50+ chapters unattended β€” a full series to video overnight. What I'd do differently
If I started over, I'd spend more time on the training data for panel detection up front instead of iterating on model tweaks β€” data diversity beat architecture fiddling every time. I'd also build the evaluation harness earlier: a small set of "golden" chapters across art styles to regression-test every change against.
Try it

Repo: github.com/aashish254/Aniflow (MIT). There's a web UI, a v0.1.0 release, and tutorial videos in the README. If you're into self-hosting, computer vision, or the manhwa scene β€” I'd love your feedback, and contributors are very welcome.

── more in #computer-vision 4 stories Β· sorted by recency
── more on @aniflow 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/how-i-built-an-ai-st…] indexed:0 read:3min 2026-09-29 Β· β€”