{"slug": "ai-models-that-make-images-video-and-speech-180-counted", "title": "AI Models That Make Images, Video and Speech: 180 Counted", "summary": "OpenRouter's public model catalog listed 180 of 625 routes, or 29 percent, producing non-text output such as images, video, speech, transcripts and embeddings as of September 25, 2026, according to a snapshot taken by the catalog's API. Image routes led with 57, followed by embeddings at 37 and video at 29, and 75 of the 180 non-text routes were added in the 90 days to September 25 versus 129 of 445 text-only routes, a 42 percent versus 29 percent pace. No usable price is printed for 39 non-text routes, including all 29 video routes, and the snapshot still listed Sora 2 Pro a day after OpenAI's scheduled removal of the Sora 2 API.", "body_md": "Most talk about AI models is about chat. But in the snapshot of OpenRouter’s public model catalog we took on September 25, 2026, 180 of the 625 routes produce something other than text: images, video, speech, transcripts and the number lists that power search. They come from 35 vendors, and they are being added faster than chat models.\n\nA route here means one model offered through the catalog’s single API, so a model with a batch or free variant counts more than once. The count below is what a developer or marketer can reach through one gateway, not a census of every model in the world. It also shows a practical gap: for 39 of these routes, including every video model, the catalog prints no usable price.\n\n1. 01180 of 625 routes, 29 percent, output something other than text.Image leads with 57 routes, then embeddings with 37 and video with 29.\n2. 02Non-text routes are arriving faster.75 of the 180 were listed in the 90 days to September 25, against 129 of 445 text-only routes: 42 percent against 29.\n3. 03No price is shown for 39 non-text routes.All 29 video routes, six rerank routes, two music routes and two image routes list zero in every price field.\n4. 04A listing is not proof a model is available.The snapshot still listed Sora 2 Pro a day after OpenAI’s scheduled removal of the Sora 2 API.\n\n## 01 — The countMore than a quarter of the catalog does not *talk*\n\n##### Routes with a non-text output\n\nCounted from every route in the all-modality catalog on September 25. The other 445 output text only.\n\n##### Vendor namespaces behind the 180 routes\n\nOpenAI has the most with 23, then Google with 20, Recraft with 16 and Qwen with 10.\n\n##### Every price field reads zero\n\nNot counting six routes marked free or OpenRouter’s two automatic routers, whose price depends on the model they pick.\n\nThe pace is the clearest signal. Non-text listings ran at 29 in July and 27 in August, and 18 more arrived in the first 25 days of September. For comparison with language models, our [census of context and output limits](https://www.digitalapplied.com/blog/model-context-window-output-limit-census) covers the text side of the market.\n\n## 02 — The tableRoutes by output type\n\nSome output types need a word of explanation. Embeddings turn text into lists of numbers so software can find similar content; rerank models sort search results by relevance. Transcription turns speech into text and speech turns text into a voice. Audio covers two Google music models and two OpenAI voice-chat models. Decisions is a newer type that returns a structured choice instead of prose; our post on [TypeSafe’s Jev](https://www.digitalapplied.com/blog/typesafe-jev-system-one-model-typed-decisions) explains it.\n\n| Source: OpenRouter models API, all output modalities, observed September 25, 2026 at 08:01 UTC. “Added” means listed on or after June 27, 2026. |  |  |  | \n|---|---|---|---|\n| Output | Routes | Added in 90 days | No price shown | \n|---|---|---|---|\n| Image | 57 | 23 | 2 | \n| Embeddings | 37 | 7 | 0 | \n| Video | 29 | 13 | 29 | \n| Transcription | 24 | 14 | 0 | \n| Speech | 20 | 13 | 0 | \n| Rerank | 7 | 3 | 6 | \n| Audio | 4 | 0 | 2 | \n| Decisions | 2 | 2 | 0 | \n| All non-text | 180 | 75 | 39 | \n| Text only (comparison) | 445 | 129 | — | \n\n#### Non-text routes by output type\n\nOpenRouter models API, observed September 25, 2026\nFifteen routes produce text as well as another output, mostly Google and OpenAI image models that can reply in words and pictures. Each counts once under its non-text type. Among speech routes, 16 of 20 list their voices: Deepgram’s Aura-2 lists 90, Kokoro 54 and MiniMax’s Speech 2.8 models 45 each, while Microsoft’s MAI Voice 2 lists four. Our [ranking of text-to-speech models](https://www.digitalapplied.com/blog/best-text-to-speech-models-september-2026-ranked-priced) compares their quality and vendor prices.\n\n## 03 — The listEvery video route in the catalog\n\nMost of the 29 video routes turn a prompt or a still image into a clip. Nine also accept video, for editing, extending or upscaling footage, and six also accept audio as an input.\n\n| Source: OpenRouter models API, observed September 25, 2026. “Listed” is the catalog’s own date, not the vendor’s release date. |  |  | \n|---|---|---|\n| Catalog ID | Listed | Accepts | \n|---|---|---|\n| black-forest-labs/flux-video-edit | Sep 10, 2026 | Text, video | \n| minimax/hailuo-3-max | Sep 2, 2026 | Text, image | \n| alibaba/wan-3.0-prime | Aug 27, 2026 | Text, image | \n| alibaba/wan-3.0 | Aug 24, 2026 | Text, image | \n| heygen/avatar-iv | Aug 24, 2026 | Text, image, audio | \n| black-forest-labs/flux-video-upscale | Aug 19, 2026 | Text, video | \n| bytedance/seedance-2.0-mini | Aug 12, 2026 | Text, image, audio, video | \n| bytedance/seedance-2.5 | Aug 7, 2026 | Text, image, audio, video | \n| black-forest-labs/flux-3-video | Aug 4, 2026 | Text, image, video | \n| minimax/hailuo-3 | Jul 29, 2026 | Text, image, audio, video | \n| runway/aleph-2 | Jul 29, 2026 | Text, image, video | \n| runway/gen-4.5 | Jul 29, 2026 | Text, image | \n| x-ai/grok-imagine-video-1.5 | Jul 20, 2026 | Text, image | \n| alibaba/happyhorse-1.1 | Jun 24, 2026 | Text, image | \n| alibaba/happyhorse-1.0 | Jun 24, 2026 | Text, image | \n| x-ai/grok-imagine-video | May 18, 2026 | Text, image | \n| kwaivgi/kling-v3.0-pro | Apr 29, 2026 | Text, image | \n| kwaivgi/kling-v3.0-std | Apr 29, 2026 | Text, image | \n| google/veo-3.1-fast | Apr 24, 2026 | Text, image | \n| google/veo-3.1-lite | Apr 23, 2026 | Text, image | \n| kwaivgi/kling-video-o1 | Apr 20, 2026 | Text, image | \n| minimax/hailuo-2.3 | Apr 20, 2026 | Text, image | \n| alibaba/wan-2.7 | Apr 15, 2026 | Text, image | \n| bytedance/seedance-2.0 | Apr 15, 2026 | Text, image, audio, video | \n| bytedance/seedance-2.0-fast | Apr 15, 2026 | Text, image, audio, video | \n| alibaba/wan-2.6 | Mar 28, 2026 | Text, image | \n| bytedance/seedance-1-5-pro | Mar 23, 2026 | Text, image | \n| openai/sora-2-pro | Mar 23, 2026 | Text, image | \n| google/veo-3.1 | Mar 23, 2026 | Text, image | \n\nOne row shows why a catalog is a starting point, not an authority. [OpenAI’s deprecations page](https://developers.openai.com/api/docs/deprecations) scheduled the removal of its Videos API and the Sora 2 models, including `sora-2-pro`, for September 24. Our snapshot a day later still listed `openai/sora-2-pro`. The catalog also shows an expiry date of November 11, 2026 on `bytedance/seedance-1-5-pro`, the only non-text route carrying one.\n\n## 04 — The catchWhat the price fields cannot tell you\n\nFor chat models, a catalog price is a cost per million tokens and can be compared directly. For other outputs it cannot. Every video route lists zero in every price field, so the catalog gives no way to budget a clip. Image routes mostly carry their cost in image fields rather than the usual input and output fields: 37 of 57 list zero for both of those.\n\nTranscription shows the unit problem most sharply. Its raw input prices run from 0.00000125 to 0.36 across 24 routes, a spread that reflects different billing units, such as tokens, seconds or minutes, rather than a real price gap. The catalog does not say which unit each route uses.\n\nA zero in a catalog price field means the catalog has no price in that field, not that the model costs nothing. Six routes in this set are explicitly marked free. For the rest, read the vendor’s own pricing page before you plan spend, and run a small paid test to see what the invoice actually counts.\n\n## 05 — MethodologyHow we counted\n\nOne snapshot, one rule per column, no inference from model names.\n\n- What was collected\n- Every route in OpenRouter’s models API with all output modalities requested: 625 records, each with its listed input and output types, listing date, price fields, voices and expiry date.\n- As-of date\n- September 25, 2026 at 08:01 UTC. Routes added or removed after that moment are not reflected. OpenAI’s deprecations page was read on September 29, 2026.\n- Classification\n- A route counts as non-text if any listed output is not text. Output type is the catalog’s own label. “Added in 90 days” uses the catalog’s listing date, on or after June 27, 2026. “No price shown” means every price field is zero, excluding routes marked free and the two automatic routers.\n- Counting rules\n- Batch, free and alias variants count as separate routes, as the catalog lists them. Vendors are counted by catalog namespace, excluding OpenRouter’s own routers and counting the ~typesafe alias under typesafe. No route listed more than one non-text output type.\n- Known limitations\n- One gateway’s catalog, not the whole market: models sold only direct by their vendors are missing. Listing dates are not release dates, and price fields were recorded raw, without converting units.\n- Refresh\n- Refreshed in place from each new catalog snapshot, with counts and the video list updated and the as-of date changed.\n\n## 06 — ConclusionChoosing a non-text model starts where the catalog stops\n\n### Use the catalog to find candidates, then confirm price, units and availability on each vendor’s own page before you build\n\nImage, video and voice models are now a large and fast-growing part of what one API key can reach, but the catalog describes them less completely than chat models. For embeddings, our [embedding cost calculator](https://www.digitalapplied.com/blog/embedding-model-cost-calculator-vendor-comparison-2026) does the conversion for you. If you want these models built into a content or product workflow with real costs attached, our [AI transformation team](https://www.digitalapplied.com/services/ai-transformation) can help.", "url": "https://wpnews.pro/news/ai-models-that-make-images-video-and-speech-180-counted", "canonical_source": "https://www.digitalapplied.com/blog/ai-models-that-make-images-video-speech-catalog-count", "published_at": "2026-09-26 00:00:00+00:00", "updated_at": "2026-09-29 04:48:49.705868+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-products", "computer-vision", "ai-infrastructure"], "entities": ["OpenRouter", "OpenAI", "Google", "Recraft", "Qwen", "Sora 2 Pro", "Deepgram", "Microsoft"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-models-that-make-images-video-and-speech-180-counted", "markdown": "https://wpnews.pro/news/ai-models-that-make-images-video-and-speech-180-counted.md", "text": "https://wpnews.pro/news/ai-models-that-make-images-video-and-speech-180-counted.txt", "jsonld": "https://wpnews.pro/news/ai-models-that-make-images-video-and-speech-180-counted.jsonld"}}