{"slug": "deepseek-v4-flash-and-qwen3-8-27b-now-available-for-telnyx-inference", "title": "DeepSeek-V4-Flash and Qwen3.8-27B Now Available for Telnyx Inference", "summary": "Telnyx has added DeepSeek-V4-Flash and Qwen3.8-27B to its Inference API, offering cost-efficient and low-latency options for high-volume and classification tasks. DeepSeek-V4-Flash is priced at $0.13 per 1M input tokens and $0.26 per 1M output tokens, while Qwen3.8-27B costs $0.40 per 1M input tokens and $3.00 per 1M output tokens. Both models run on Telnyx-owned GPU infrastructure and are accessible via the existing OpenAI-compatible endpoint without SDK updates.", "body_md": "[Contact us](https://telnyx.com/contact-us)\n\n[Log in](https://portal.telnyx.com)\n\nTwo new smaller, faster LLMs are now available on the [Telnyx Inference API](https://telnyx.com/products/inference): DeepSeek-V4-Flash for cost-efficient high-volume workloads, and Qwen3.8-27B for low-latency routing and classification tasks. Both run on Telnyx-owned GPU infrastructure through the existing OpenAI-compatible Chat Completions endpoint, with no new endpoints or SDK updates required.\n\n`deepseek-ai/DeepSeek-V4-Flash-0731`\n\n. A smaller, faster DeepSeek variant tuned for cost-efficient inference at high volume.`Qwen/Qwen3.8-27B`\n\n. A 27B-parameter model from the Qwen team, sized for low-latency tasks where a full-scale LLM is overkill.| Token Type | DeepSeek-V4-Flash | Qwen3.8-27B |\n|---|---|---|\n| Input | $0.13 / 1M tokens | $0.40 / 1M tokens |\n| Cached Input | $0.03 / 1M tokens | $0.05 / 1M tokens |\n| Output | $0.26 / 1M tokens | $3.00 / 1M tokens |\n\nFull rate details on the [inference pricing page](https://telnyx.com/pricing/inference-api).\n\n`deepseek-ai/DeepSeek-V4-Flash-0731`\n\nor `Qwen/Qwen3.8-27B`\n\nfrom the model dropdown.\n\n```\ncurl https://api.telnyx.com/v2/ai/chat/completions   -H \"Authorization: Bearer $TELNYX_API_KEY\"   -H \"Content-Type: application/json\"   -d '{\n    \"model\": \"deepseek-ai/DeepSeek-V4-Flash-0731\",\n    \"messages\": [\n      {\"role\": \"user\", \"content\": \"Classify this support ticket into one of: billing, technical, account.\"}\n    ]\n  }'\n```\n\n**Learn more** in the [Inference API docs](https://developers.telnyx.com/docs/inference/models) or on the [pricing page](https://telnyx.com/pricing/inference-api).", "url": "https://wpnews.pro/news/deepseek-v4-flash-and-qwen3-8-27b-now-available-for-telnyx-inference", "canonical_source": "https://telnyx.com/release-notes/deepseek-v4-flash-qwen3-8-27b-inference", "published_at": "2026-08-21 12:00:00+00:00", "updated_at": "2026-08-21 22:42:34.184171+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-infrastructure"], "entities": ["Telnyx", "DeepSeek-V4-Flash", "Qwen3.8-27B", "Qwen team", "Telnyx Inference API"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-flash-and-qwen3-8-27b-now-available-for-telnyx-inference", "markdown": "https://wpnews.pro/news/deepseek-v4-flash-and-qwen3-8-27b-now-available-for-telnyx-inference.md", "text": "https://wpnews.pro/news/deepseek-v4-flash-and-qwen3-8-27b-now-available-for-telnyx-inference.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-flash-and-qwen3-8-27b-now-available-for-telnyx-inference.jsonld"}}