{"slug": "keeping-one-character-consistent-across-a-whole-article-s-ai-illustrations", "title": "Keeping One Character Consistent Across a Whole Article's AI Illustrations", "summary": "A developer built InkDoo, a tool that generates a consistent set of hand-drawn explainer illustrations for an article, all featuring the same recurring character. The system uses a planner LLM to select ideas and scene descriptions, an image model driven by a strict prompt template with a fixed \"visual DNA\" block and a literal character spec, and an exporter that inserts hosted image URLs into the Markdown. The developer reports that keeping the variable prompt surface small and requiring the character to perform each image's core conceptual action were the biggest factors in visual consistency.", "body_md": "If you've tried to illustrate a long blog post with an image model, you know how it goes. You ask for five pictures and get five different art styles, three versions of your \"mascot\", and a stock-photo handshake you never asked for. Each image is fine by itself. Put them together in one post, though, and it looks like a ransom note.\n\nI build [InkDoo](https://inkdoo.app?utm_source=devto), a tool that takes an article and returns a full set of hand-drawn explainer illustrations, all starring the same character. This post covers the parts that actually made the set look consistent. Spoiler: most of it is boring prompt plumbing, not model magic.\n\nArticle in. A planner LLM reads the text and returns JSON: which ideas are worth a picture, which paragraph each picture goes after, and a scene description for each. Then an image model draws every scene from a strict prompt template. Finally, an exporter drops the hosted image URLs back into your Markdown at the right spots. Three stages, and the consistency work happens in all three.\n\nThe biggest single win was writing each character as a short, very literal visual spec, then pasting that exact string into *every* image prompt. Here is the default character, Mochi:\n\n```\n\"Mochi\", a round low squishy cat-like blob drawn ONLY as a hollow black\noutline (white inside, never filled black): a wide soft dumpling-shaped body\nsitting flat, two tiny triangle ears, two short sleepy horizontal line eyes,\na tiny 'w' mouth, small nub paws, one thin curly tail.\n```\n\nA few things I learned writing these:\n\nInkDoo ships six of these (Mochi, Doo, Blot, Stub, Folio, Puff), and they all follow the same spec format.\n\nA consistent character that just stands in the corner waving is decoration. The planner's system prompt has one rule that changed the output more than anything else:\n\n```\n${N} must PERFORM the core conceptual action of each image (pulling, carrying,\nsieving, weighing, stitching, guarding, pushing, folding, unpacking...).\nIf the image still works without ${N}, ${N} is too decorative â€” rewrite.\n```\n\nBecause the character is part of the metaphor, every image has to draw it at a size and angle where its features are readable. A tiny figure in the background is where likeness falls apart, so this rule helps consistency too.\n\nEvery image prompt has the same fixed \"visual DNA\" block: pure white background, minimalist black hand-drawn wobbly line art, at least 35% empty space, and sparse handwritten notes in only three colors. Each color has a job: orange for the main flow and arrows, red for the key warning or result, blue for side notes. It also carries a list of things to avoid: gradients, shadows, PPT infographics, cute-poster energy.\n\nOnly a few fields change from image to image: theme, structure type, core idea, composition, and two to five short labels. The planner picks the structure from a small, fixed list (Workflow, Before/after, Concept metaphor, Route map, Mini comicâ€¦). It also has to invent a fresh, low-tech physical metaphor using one or two objects: a funnel, a scale, a drawer, a broken machine.\n\nKeeping the variable surface small is the real trick. The less each prompt is allowed to vary, the more the set reads as one series.\n\n(Credit where it's due: the illustration method, meaning one idea per image, white background and sparse colored annotations, is adapted from the MIT-licensed [Ian Xiaohei Illustrations](https://github.com/helloianneo/ian-xiaohei-illustrations) project. The characters are original.)\n\nUsers can upload their own character as a reference picture. That creates an asymmetry. The image call can take the reference (InkDoo sends it to the edits endpoint), but the planner LLM is text-only. So a custom character gets two descriptions:\n\n`desc`, for the image model: \"the user's own original character, exactly as shown in the attached reference imageâ€¦ keep its silhouette, proportions, face and distinctive featuresâ€¦ but redraw it in this minimalist black hand-drawn line-art styleâ€¦ ignore the reference background.\"`planDesc`, for the planner: a text-only version built from the name and any look description the user typed.\nUser text is sanitized (control characters, quotes and braces stripped, length capped) before it gets near a prompt. If the custom character is empty, the resolver quietly falls back to Mochi instead of producing a character-less set.\n\nPlanning is free in InkDoo, and generating images costs credits. So I didn't want people re-planning just because they picked a different character halfway through. The planner writes scenes in English with the character's name in them, so switching is a careful, Unicode-aware whole-word replace of the old name with the new one. Mochi's scene (\"Mochi sits on the overflowing inbox\") becomes Doo's scene, with no extra LLM call.\n\nThe planner returns an `after` anchor for each shot: the first ~20 characters of the paragraph the image should follow, copied verbatim. The exporter splits the article into blocks, normalizes whitespace, and matches each anchor to a block. Then it writes out Markdown (or WeChat-friendly HTML) with the hosted image URLs in place. Matching on a short prefix holds up much better than asking the model for paragraph numbers, which it miscounts.\n\nIf you want a consistent character across many AI images:\n\nIf you'd rather not build the plumbing yourself, you can try the whole pipeline at [inkdoo.app](https://inkdoo.app?utm_source=devto). Paste a post or drop in a link or a .md/.docx file. New accounts get two free images. I'd love to hear how you're handling character consistency in your own projects.", "url": "https://wpnews.pro/news/keeping-one-character-consistent-across-a-whole-article-s-ai-illustrations", "canonical_source": "https://dev.to/xianyu110/keeping-one-character-consistent-across-a-whole-articles-ai-illustrations-1p5l", "published_at": "2026-10-06 04:02:08+00:00", "updated_at": "2026-10-06 04:18:03.724843+00:00", "lang": "en", "topics": ["ai-tools", "generative-ai", "large-language-models", "ai-products"], "entities": ["InkDoo", "Mochi", "Ian Xiaohei Illustrations"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/keeping-one-character-consistent-across-a-whole-article-s-ai-illustrations", "markdown": "https://wpnews.pro/news/keeping-one-character-consistent-across-a-whole-article-s-ai-illustrations.md", "text": "https://wpnews.pro/news/keeping-one-character-consistent-across-a-whole-article-s-ai-illustrations.txt", "jsonld": "https://wpnews.pro/news/keeping-one-character-consistent-across-a-whole-article-s-ai-illustrations.jsonld"}}