{"slug": "show-hn-1endpoint-cheaper-access-to-ai-models", "title": "Show HN: 1endpoint – Cheaper access to AI models", "summary": "1endpoint, a new API gateway, offers access to multiple AI models through a single endpoint with usage-based pricing, starting at $0.0420 per 1M input tokens for GLM 5.2. The service includes prompt caching, spend tracking, and claims never to downgrade models, allowing developers to switch models without rewriting integrations.", "body_md": "# Multiple models. One endpoint.\n\nUse one compatible API, switch models without rewriting your integration, and pay for the tokens you actually use.\n\n/api/v1/chat/completions\n\n- Prompt Caching\n- Spend Tracking\n- Never Downgraded\n\n`glm-5.2`\n\n**$0.0420**\n\n*USD per 1M input tokens*\n\n`gpt-5.6-luna`\n\n**$0.0500**\n\n*USD per 1M input tokens*\n\n`glm-5.3-flash`\n\n**$0.0555**\n\n*USD per 1M input tokens*\n\n`deepseek-v4-flash`\n\n**$0.0660**\n\n*USD per 1M input tokens*\n\n`deepseek-v4-flash-vision-exp`\n\n**$0.0660**\n\n*USD per 1M input tokens*\n\n`minimax-m3`\n\n**$0.0900**\n\n*USD per 1M input tokens*\n\n`gpt-5.6-terra`\n\n**$0.1200**\n\n*USD per 1M input tokens*\n\n`glm-5.3`\n\n**$0.1260**\n\n*USD per 1M input tokens*\n\n`gemini-3.7-flash`\n\n**$0.1350**\n\n*USD per 1M input tokens*\n\n`qwen3.8-2.4t-a95b`\n\n**$0.1400**\n\n*USD per 1M input tokens*\n\n`kimi-k3`\n\n**$0.1500**\n\n*USD per 1M input tokens*\n\n`deepseek-v4-pro`\n\n**$0.1980**\n\n*USD per 1M input tokens*\n\n`gpt-5.6-sol`\n\n**$0.2000**\n\n*USD per 1M input tokens*\n\n`sonnet-5`\n\n**$0.3000**\n\n*USD per 1M input tokens*\n\n`grok-4.6`\n\n**$0.4000**\n\n*USD per 1M input tokens*\n\n`opus-5`\n\n**$0.4500**\n\n*USD per 1M input tokens*\n\n`fable-5`\n\n**$1.5000**\n\n*USD per 1M input tokens*\n\n## Compare the numbers.\n\n| Relative | |||||\n|---|---|---|---|---|---|\n| GLM 5.2 | 0.0420 | 0.0078 | 0.1320 | 0.05106 | |\n| GPT 5.6 Luna | 0.0500 | 0.0050 | 0.3000 | 0.0935 | |\n| GLM 5.3 Flash | 0.0555 | 0.0111 | 0.1850 | 0.07067 | |\n| DeepSeek V4 Flash 0731 | 0.0660 | 0.0066 | 0.1980 | 0.07392 | |\n| DeepSeek V4 Flash Vision Exp | 0.0660 | 0.0066 | 0.1980 | 0.07392 | |\n\nBlended assumes 70% of input served from cache and output equal to 25% of input volume — a reference mix for comparison only, not a billed rate. The gateway records usage after a successful response; rates shown are the currently configured rates.\n\n## Change one base URL.\n\nKeep the request shape your application already understands. Change the model ID when the workload changes.\n\n**1endpoint**\n\n`one base URL`\n\n**Chat Completions**`POST /chat/completions`\n\n**Responses**`POST /responses`\n\n**Messages**`POST /messages`\n\n## Cache decides the bill.\n\nMessage 12 of a conversation costs 3.7× less than it would without cache — because every message resends everything before it, and we charge full price only for the part that is new.\n\nGLM 5.2 rates, USD per 1M tokens — a cache hit is billed at 5× less than a miss. Illustration: a conversation of 12 messages, each one request that adds 10,000 input and 1,000 output tokens and resends everything before it, priced per 1,000 requests. Token counts are illustrative; the rates are published. Cache ratio differs per model.\n\n## Spend less on every token.\n\nPay only for what you use. Input starts at $0.0420 per 1M tokens — lower cache rates are applied automatically, per model.\n\nUsage-based pricing · 1,000 credits = $1", "url": "https://wpnews.pro/news/show-hn-1endpoint-cheaper-access-to-ai-models", "canonical_source": "https://1endpoint.dev", "published_at": "2026-08-30 11:13:32+00:00", "updated_at": "2026-08-30 11:21:50.005050+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-tools", "developer-tools"], "entities": ["1endpoint", "GLM 5.2", "GPT 5.6 Luna", "DeepSeek V4 Flash", "Gemini 3.7 Flash"], "alternates": {"html": "https://wpnews.pro/news/show-hn-1endpoint-cheaper-access-to-ai-models", "markdown": "https://wpnews.pro/news/show-hn-1endpoint-cheaper-access-to-ai-models.md", "text": "https://wpnews.pro/news/show-hn-1endpoint-cheaper-access-to-ai-models.txt", "jsonld": "https://wpnews.pro/news/show-hn-1endpoint-cheaper-access-to-ai-models.jsonld"}}