{"slug": "tinker-glm-5-3-fine-tuning", "title": "Tinker: GLM 5.3 Fine-Tuning", "summary": "Tinker, an AI model fine-tuning platform, has added GLM 5.3 (256K) to its training lineup, alongside other models such as Inkling, Nemotron-3.5-Lightning-30B-A3B, and Qwen3.8-27B, with prices per million tokens and an 80% discount on cached prefill tokens. The platform also offers serverless inference in beta for Inkling and Inkling-Small, with checkpoint storage charged at $0.10 per GB per month.", "body_md": "# Models & Pricing\n\nAll prices are per million tokens. Checkpoint storage is charged at $0.10 per GB per month.\n\n## Training\n\nWe provide an 80% discount on cached prefill tokens.\n\n| Model | Tinker ID | Context | Size | Arch | Type | PrefillCached: 80% discount | Sample | Train |\n|---|---|---|---|---|---|---|---|---|\n|\n\n[Inkling (256K)](https://huggingface.co/thinkingmachines/Inkling)Limited-time 50% discount[Inkling-Small](https://huggingface.co/thinkingmachines/Inkling-Small)Limited-time 50% discount[Inkling-Small (256K)](https://huggingface.co/thinkingmachines/Inkling-Small)Limited-time 50% discount[Nemotron-3.5-Lightning-30B-A3B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16)Limited-time 50% discount[Nemotron-3.5-Lightning-30B-A3B (256K)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16)Limited-time 50% discount[Nemotron-3-Ultra-550B-A55B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16)Limited-time 50% discount[Nemotron-3-Ultra-550B-A55B (256K)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16)Limited-time 50% discount[Nemotron-3-Super-120B-A12B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16)Limited-time 50% discount[Nemotron-3-Super-120B-A12B (256K)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16)Limited-time 50% discount[Nemotron-3-Nano-30B-A3B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16)Limited-time 50% discount[GLM-5.3 (256K)](https://huggingface.co/zai-org/GLM-5.3)[Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6)[Kimi-K2.6 (128K)](https://huggingface.co/moonshotai/Kimi-K2.6)[Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)[Qwen3.8-27B (256K)](https://huggingface.co/Qwen/Qwen3.8-27B)[Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B)[Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B)[Retiring September 2](../model-deprecations/)[Qwen3.5-397B-A17B](https://huggingface.co/Qwen/Qwen3.5-397B-A17B)[Qwen3.5-397B-A17B (256K)](https://huggingface.co/Qwen/Qwen3.5-397B-A17B)[Qwen3.5-35B-A3B-Base](https://huggingface.co/Qwen/Qwen3.5-35B-A3B-Base)[Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)[Qwen3.5-9B-Base](https://huggingface.co/Qwen/Qwen3.5-9B-Base)[Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)[Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)[GPT-OSS-120B](https://huggingface.co/openai/gpt-oss-120b)[GPT-OSS-120B (128K)](https://huggingface.co/openai/gpt-oss-120b)[GPT-OSS-20B](https://huggingface.co/openai/gpt-oss-20b)[DeepSeek-V3.1](https://huggingface.co/deepseek-ai/DeepSeek-V3.1)## Serverless Inference (Beta)\n\nServerless inference is currently in beta and available for Inkling and Inkling-Small only. We do not recommend it for intensive production use until it is out of beta. If you're interested in production use, email us at [ [email protected]](/cdn-cgi/l/email-protection#98ecf1f6f3fdead8ecf0f1f6f3f1f6fff5f9fbf0f1f6fdebb6f9f1a7ebedfaf2fdfbeca5cbfdeaeefdeaf4fdebebbdaaa8d1f6fefdeafdf6fbfdbdaaa8eff9f1ecf4f1ebecbef9f5e8a3faf7fce1a5cff0f1fbf0bdaaa8f7eafff9f6f1e2f9ecf1f7f6bdaaa8e1f7edbdaaa8eafde8eafdebfdf6ecbdabd9bda8d9bda8d9cff0f1fbf0bdaaa8f5f7fcfdf4ebbdaaa8e1f7edbdaaa8f6fdfdfcbdabd9bda8d9bda8d9dde0e8fdfbecfdfcbdaaa8eef7f4edf5fdbdabd9bda8d9bda8d9d4f9ecfdf6fbe1bdaaa8eafde9edf1eafdf5fdf6ecebbdabd9bda8d9bda8d9cdebfdbdaaa8fbf9ebfdbdabd9bda8d9bda8d9) to join the waitlist. Please include which models you need, your expected volume and latency requirements, and your use case.\n\n| Model | Tinker ID | Context | Prefill (Input) | Sample (Output) |\n|---|---|---|---|---|\n|\n\n[Inkling](https://huggingface.co/thinkingmachines/Inkling)## Pricing Terms\n\n**Prefill**: Processing input/prompt tokens (forward pass only)** Cached prefill**: The smaller price under each prefill price; applies to input tokens that hit the prompt cache (80% off)** Sample**: Generating output tokens (forward pass + sampling)** Train**: Forward and backward pass for gradient computation** Context**: Maximum sequence length. Models with`:peft:`\n\nsuffix support extended context at higher prices.**Tinker ID**: The exact string to pass to`create_lora_training_client(base_model=...)`\n\nor`create_sampling_client(base_model=...)`\n\nMoE models are priced by active parameters, making them significantly more cost-effective than dense models of similar quality.\n\n## Model Types\n\n**Base**: Raw pretrained models with no chat or instruction tuning. Best for post-training research or running the full post-training pipeline yourself.**Reasoning**: Always produce chain-of-thought before their answer. Highest intelligence, higher latency and token cost.** Hybrid**: Run in both thinking and non-thinking modes. They reason by default, but chain-of-thought can be disabled via a renderer or argument for faster, cheaper direct answers.**Vision**: Vision-language models that accept images alongside text. Shown as a`+ Vision`\n\nsuffix on the underlying type (for example,`Hybrid + Vision`\n\n).**Audio**: Models that accept audio alongside text. Shown as a`+ Audio`\n\nsuffix on the underlying type.\n\n**Architecture** is either **Dense** (all parameters active per token) or **MoE** (mixture-of-experts, only a subset of parameters active per token). MoE models are highlighted in amber.\n\n## Choosing a Model\n\n**Cost-effective**: Use MoE models (highlighted in amber)** Research/post-training**: Use Base models** Task-specific fine-tuning**: Start with a Hybrid model** Low latency**: Use a Hybrid model with chain-of-thought disabled** High intelligence**: Use Reasoning or Hybrid models (chain-of-thought)** Vision tasks**: Use models with Vision in the type\n\n## Machine-Readable Pricing\n\nIf you want to use this data programmatically (cost estimation, model pickers, dashboards), don't scrape the tables above. The data behind them is published as JSON alongside this page, and those files are the stable interface for scripts (the tables' HTML is presentational and may change):\n\n[models.json](../models.json): the Training table. One object per model with`name`\n\n,`tinker_id`\n\n,`context`\n\n,`size`\n\n,`arch`\n\n,`type`\n\n,`url`\n\n, and per-million-token prices`prefill`\n\n,`cached_prefill`\n\n,`sample`\n\n,`train`\n\n(strings like`\"$0.374\"`\n\n). Temporarily discounted models also carry`original_*`\n\nprice fields and a`note`\n\n.[serverless.json](../serverless.json): the Serverless Inference table, with`name`\n\n,`tinker_id`\n\n,`context`\n\n,`url`\n\n,`input`\n\n,`cached_input`\n\n,`output`\n\n.\n\nFor example:\n\n``` python\nimport json, urllib.request\n\nurl = \"https://tinker-docs.thinkingmachines.ai/tinker/models.json\"\nmodels = json.load(urllib.request.urlopen(url))\nfor m in models:\n    print(m[\"tinker_id\"], m[\"train\"])\n```\n\n## Retired Models\n\nThese models have been retired and can no longer be used for training or inference, grouped by retirement date. See [Model deprecations](../model-deprecations/) for the recommended replacement for each.\n\n### July 12, 2026\n\n**Kimi:**`Kimi-K2.5`\n\n### June 12, 2026\n\n**Qwen:**`Qwen3-235B-A22B-Instruct-2507`\n\n,`Qwen3-VL-235B-A22B-Instruct`\n\n,`Qwen3.5-35B-A3B`\n\n,`Qwen3.5-27B`\n\n,`Qwen3-32B`\n\n,`Qwen3-30B-A3B`\n\n,`Qwen3-30B-A3B-Instruct-2507`\n\n,`Qwen3-VL-30B-A3B-Instruct`\n\n,`Qwen3-30B-A3B-Base`\n\n,`Qwen3-8B-Base`\n\n,`Qwen3-4B-Instruct-2507`\n\n**Llama:**`Llama-3.3-70B-Instruct`\n\n,`Llama-3.1-70B`\n\n,`Llama-3.1-8B`\n\n,`Llama-3.1-8B-Instruct`\n\n,`Llama-3.2-3B`\n\n,`Llama-3.2-1B`\n\n**DeepSeek:**`DeepSeek-V3.1-Base`\n\n**Kimi:**`Kimi-K2-Thinking`", "url": "https://wpnews.pro/news/tinker-glm-5-3-fine-tuning", "canonical_source": "https://tinker-docs.thinkingmachines.ai/tinker/models/", "published_at": "2026-08-29 10:12:48+00:00", "updated_at": "2026-08-29 10:48:55.059810+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-tools", "large-language-models"], "entities": ["Tinker", "GLM 5.3", "Inkling", "Nemotron-3.5-Lightning-30B-A3B", "Qwen3.8-27B", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/tinker-glm-5-3-fine-tuning", "markdown": "https://wpnews.pro/news/tinker-glm-5-3-fine-tuning.md", "text": "https://wpnews.pro/news/tinker-glm-5-3-fine-tuning.txt", "jsonld": "https://wpnews.pro/news/tinker-glm-5-3-fine-tuning.jsonld"}}