cd /news/ai-infrastructure/show-hn-1endpoint-cheaper-access-to-… · home topics ai-infrastructure article
[ARTICLE · art-115669] src=1endpoint.dev ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Show HN: 1endpoint – Cheaper access to AI models

1endpoint, a new API gateway, offers access to multiple AI models through a single endpoint with usage-based pricing, starting at $0.0420 per 1M input tokens for GLM 5.2. The service includes prompt caching, spend tracking, and claims never to downgrade models, allowing developers to switch models without rewriting integrations.

read2 min views6 publishedAug 30, 2026
Show HN: 1endpoint – Cheaper access to AI models
Image: source

Use one compatible API, switch models without rewriting your integration, and pay for the tokens you actually use. /api/v1/chat/completions

  • Prompt Caching
  • Spend Tracking
  • Never Downgraded

glm-5.2

$0.0420

USD per 1M input tokens

gpt-5.6-luna $0.0500

USD per 1M input tokens

glm-5.3-flash $0.0555

USD per 1M input tokens

deepseek-v4-flash $0.0660

USD per 1M input tokens

deepseek-v4-flash-vision-exp $0.0660

USD per 1M input tokens

minimax-m3

$0.0900

USD per 1M input tokens

gpt-5.6-terra $0.1200

USD per 1M input tokens

glm-5.3

$0.1260

USD per 1M input tokens

gemini-3.7-flash $0.1350

USD per 1M input tokens

qwen3.8-2.4t-a95b $0.1400

USD per 1M input tokens

kimi-k3

$0.1500

USD per 1M input tokens

deepseek-v4-pro $0.1980

USD per 1M input tokens

gpt-5.6-sol $0.2000

USD per 1M input tokens

sonnet-5

$0.3000

USD per 1M input tokens

grok-4.6

$0.4000

USD per 1M input tokens

opus-5

$0.4500

USD per 1M input tokens

fable-5

$1.5000

USD per 1M input tokens

Compare the numbers. #

Relative
GLM 5.2 0.0420 0.0078 0.1320 0.05106
GPT 5.6 Luna 0.0500 0.0050 0.3000 0.0935
GLM 5.3 Flash 0.0555 0.0111 0.1850 0.07067
DeepSeek V4 Flash 0731 0.0660 0.0066 0.1980 0.07392
DeepSeek V4 Flash Vision Exp 0.0660 0.0066 0.1980 0.07392

Blended assumes 70% of input served from cache and output equal to 25% of input volume — a reference mix for comparison only, not a billed rate. The gateway records usage after a successful response; rates shown are the currently configured rates.

Change one base URL. #

Keep the request shape your application already understands. Change the model ID when the workload changes.

1endpoint

one base URL

Chat CompletionsPOST /chat/completions

ResponsesPOST /responses

MessagesPOST /messages

Cache decides the bill. #

Message 12 of a conversation costs 3.7× less than it would without cache — because every message resends everything before it, and we charge full price only for the part that is new.

GLM 5.2 rates, USD per 1M tokens — a cache hit is billed at 5× less than a miss. Illustration: a conversation of 12 messages, each one request that adds 10,000 input and 1,000 output tokens and resends everything before it, priced per 1,000 requests. Token counts are illustrative; the rates are published. Cache ratio differs per model.

Spend less on every token. #

Pay only for what you use. Input starts at $0.0420 per 1M tokens — lower cache rates are applied automatically, per model.

Usage-based pricing · 1,000 credits = $1

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @1endpoint 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-1endpoint-ch…] indexed:0 read:2min 2026-08-30 ·