{"slug": "the-fastest-and-cheapest-glm-5-3-flash-endpoint", "title": "The fastest and cheapest GLM 5.3 Flash endpoint", "summary": "RunInfra is offering GLM 5.3 Flash, an LLM from zai-org, at $0.10 per 1M input tokens and $0.40 per 1M output tokens, with a 1,048,576-token context window and OpenAI-compatible chat completions. The endpoint is listed in RunInfra Model APIs and requires a RUNINFRA_GATEWAY_KEY for access.", "body_md": "`zai-org/GLM-5.3-Flash`\n\nGLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.10 per 1M input tokens and $0.40 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions.\n\nUSD, pay per token\n\nConfirm how your client reaches this model.\n\nCheck the limits your workload must fit.\n\nSee which request modes the API supports.\n\nSet RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.\n\nVerify the company and operating credentials behind this API.\n\n© 2026 RunInfra. All rights reserved.", "url": "https://wpnews.pro/news/the-fastest-and-cheapest-glm-5-3-flash-endpoint", "canonical_source": "https://runinfra.ai/inference-api/glm-5-3-flash", "published_at": "2026-08-26 20:30:40+00:00", "updated_at": "2026-08-26 20:49:00.631834+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-infrastructure"], "entities": ["RunInfra", "zai-org", "GLM 5.3 Flash"], "alternates": {"html": "https://wpnews.pro/news/the-fastest-and-cheapest-glm-5-3-flash-endpoint", "markdown": "https://wpnews.pro/news/the-fastest-and-cheapest-glm-5-3-flash-endpoint.md", "text": "https://wpnews.pro/news/the-fastest-and-cheapest-glm-5-3-flash-endpoint.txt", "jsonld": "https://wpnews.pro/news/the-fastest-and-cheapest-glm-5-3-flash-endpoint.jsonld"}}