{"slug": "show-hn-qwen3-8-27b-api-140-tok-s-on-one-gpu", "title": "Show HN: Qwen3.8-27B API, 140 tok/s on one GPU", "summary": "Tiyuvta AI launched a hosted API for the Qwen3.8-27B model, priced at $0.38 per million input tokens, $0.20 per million cached tokens, and $2.60 per million output tokens, with a throughput of 140 tokens per second on a single GPU. New accounts receive $10.00 in free credit, with 93 of 100 places still available.", "body_md": "account\n\n# Sign in or create an account.\n\nOne step. The same screen creates the account, so there is nothing to confirm afterwards and no waiting list.\n\n**free credit**$10.00 of credit is added to your account on sign-up. 93 of 100 places left.\n\n- Base URL\n`https://api.tiyuvta.ai/v1`\n\n- Models\n`qwen/qwen3.8-27b`\n\n· Step-3.7-Flash in bring-up- Price\n- $0.38 input · $0.20 cached · $2.60 output per million tokens\n- Billing\n- Prepaid credit. At zero balance the API returns\n`402`\n\n.\n\nor\n\nBy continuing you accept the [terms](/legal/terms) and the\n[privacy terms](/legal/privacy).", "url": "https://wpnews.pro/news/show-hn-qwen3-8-27b-api-140-tok-s-on-one-gpu", "canonical_source": "https://inference.tiyuvta.ai/login?next=%2Fapp", "published_at": "2026-08-16 14:42:59+00:00", "updated_at": "2026-08-16 15:11:02.803374+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-infrastructure"], "entities": ["Tiyuvta AI", "Qwen3.8-27B"], "alternates": {"html": "https://wpnews.pro/news/show-hn-qwen3-8-27b-api-140-tok-s-on-one-gpu", "markdown": "https://wpnews.pro/news/show-hn-qwen3-8-27b-api-140-tok-s-on-one-gpu.md", "text": "https://wpnews.pro/news/show-hn-qwen3-8-27b-api-140-tok-s-on-one-gpu.txt", "jsonld": "https://wpnews.pro/news/show-hn-qwen3-8-27b-api-140-tok-s-on-one-gpu.jsonld"}}