{"slug": "iquest-q1-how-to-self-host-the-320b-agentic-coding-model", "title": "IQuest-Q1: How to Self-Host the 320B Agentic Coding Model", "summary": "IQuest released IQuest-Q1, an open-weight 320B-parameter Mixture-of-Experts coding model that activates only 15B parameters per token and supports a 524,288-token context window, with official deployment through SGLang and vLLM at tensor-parallel size 8. The model ships as safetensors across 175 shards on Hugging Face, where it has 128 likes and 778 downloads, and integrates with Claude Code and Codex CLI via custom tool-call and reasoning parsers. IQuest states the model is text-only, early-stage, and prone to agentic failure modes such as repeating failed attempts or missing constraints during real-world CLI tasks.", "body_md": "# IQuest-Q1: How to Self-Host the 320B Agentic Coding Model\n\nIQuest-Q1 is a 320B MoE coding model with 15B active params and 512K context. Here's what it takes to run it with SGLang or vLLM.\n\n## What is IQuest-Q1?\n\nIQuest-Q1 is an open-weight Mixture-of-Experts model built by IQuest for agentic coding, reasoning, and multi-step tool use. It has roughly 320 billion total parameters, but only about 15 billion are active per token thanks to its MoE routing (256 experts total, 8 activated per forward pass). It supports a context window of 524,288 tokens and ships with multi-token prediction (MTP) layers for speculative decoding. The model is distributed as safetensors across 175 shards on Hugging Face, where it has already picked up 128 likes and 778 downloads since release.\n\n## TL;DR\n\n- **IQuest-Q1** is a 320B-parameter MoE model with only 15B parameters activated per token, which keeps inference cost closer to a mid-size dense model than its total parameter count suggests.\n- The model supports a **512K token context window** , which IQuest extends to a “1m” client-side label in Claude Code without actually changing the underlying 524,288-token limit.\n- Deployment is officially supported through **SGLang and vLLM** , each with a prebuilt CUDA 13.0 Docker image and tensor-parallel size of 8, meaning you need an 8-GPU node to run it as documented.\n- The architecture uses a **hybrid attention pattern** (3 sliding-window attention layers for every 1 full attention layer) plus partial RoPE, which is a common trick for cutting memory and compute costs on long-context models.\n- IQuest-Q1 includes **recursive MTP** (2 independent layers in training, collapsed to 1 recursive layer running 8 times at inference) for speculative decoding speedups, configurable via EAGLE in both SGLang and vLLM.\n- The model integrates directly with **Claude Code and Codex CLI** through custom tool-call and reasoning parsers, positioning it as a drop-in replacement for Anthropic or OpenAI models in agentic coding harnesses.\n- IQuest is upfront about limitations: it’s **text-only** , still early-stage, and prone to the usual agentic failure modes like repeating failed attempts or missing constraints during real-world CLI tasks.\n\n## Other agents start typing. Remy starts asking.\n\nScoping, trade-offs, edge cases — the real work. Before a line of code.\n\n## What are the model’s architecture specs?\n\nIQuest-Q1 runs 88 transformer layers with a hidden dimension of 3,072. Attention uses 48 query heads and 8 key/value heads at a head dimension of 128, a grouped setup that reduces KV cache size compared to full multi-head attention. The hybrid attention pattern alternates three sliding-window attention (SWA) layers, each with a 4,096-token window, for every one full attention (FA) layer. This pattern is how the model handles a 512K context without the quadratic cost of full attention at every layer. Partial RoPE is applied across only 32 dimensions, another detail aimed at keeping long-context inference manageable.\n\nThe MoE layer selects 8 experts out of 256 per token. Combined with the 15B active parameter count, this is what makes a 320B model plausible to run at reasonable throughput, since compute scales with active parameters, not the full weight count. Memory for weights is a different story: you still need to hold all 320B parameters (plus KV cache) in GPU memory regardless of how many are active per token, which is the real constraint when planning hardware.\n\n## What hardware do you need to run it?\n\nIQuest’s own deployment commands specify `--tp-size 8` for SGLang and `--tensor-parallel-size 8` for vLLM, which means the reference deployment assumes an 8-GPU node. At 320B total parameters in bfloat16, raw weight storage alone runs into the hundreds of gigabytes before accounting for KV cache, activations, or the separate MTP draft model. This isn’t a model you run on a single consumer GPU or even a single high-end workstation card. It’s built for multi-GPU inference servers, the kind of setup typically found in datacenter or cloud GPU instances with high-bandwidth interconnects between cards.\n\nIf you don’t have access to 8 GPUs with enough combined memory, the realistic options are renting cloud GPU capacity, waiting for the community to produce quantized versions, or running it at reduced precision if and when such builds appear. The model card itself only documents the full bfloat16, tensor-parallel-8 deployment path.\n\n## How do you deploy IQuest-Q1 with SGLang or vLLM?\n\nIQuest publishes prebuilt Docker images for both serving frameworks, which removes most of the dependency wrangling that normally comes with running a brand-new model architecture.\n\nFor SGLang, pull `iquestlabworkspace/sglang-iquest-q1:cu130`, then launch the server pointing at the downloaded model weights:\n\n```\ndocker pull iquestlabworkspace/sglang-iquest-q1:cu130\n\nMODEL_ROOT=\"$(hf download IQuestLab/IQuest-Q1 --quiet)\" && \\\npython -u -m sglang.launch_server \\\n  --model-path \"$MODEL_ROOT\" \\\n  --served-model-name IQuest-Q1 \\\n  --tp-size 8 \\\n  --dtype bfloat16 \\\n  --attention-backend fa3 \\\n  --mem-fraction-static 0.85 \\\n  --tool-call-parser iquest_q1 \\\n  --reasoning-parser iquest_q1 \\\n  --enable-torch-compile \\\n  --load-format fastsafetensors\n```\n\nNote the custom `iquest_q1` tool-call and reasoning parsers. These aren’t optional extras: IQuest-Q1 uses its own chat template format for tool calls and reasoning traces, so a generic OpenAI-style parser won’t correctly extract structured output.\n\nFor vLLM, the equivalent image is `iquestlabworkspace/vllm-iquest-q1:cu130`:\n\n```\ndocker pull iquestlabworkspace/vllm-iquest-q1:cu130\n\nMODEL_ROOT=\"$(hf download IQuestLab/IQuest-Q1 --quiet)\" && \\\nvllm serve \"$MODEL_ROOT\" \\\n  --served-model-name IQuest-Q1 \\\n  --tensor-parallel-size 8 \\\n  --reasoning-parser iquest_q1 \\\n  --enable-auto-tool-choice \\\n  --tool-call-parser iquest_q1\n```\n\nBoth frameworks also support an MTP-enabled launch path, adding EAGLE-style speculative decoding flags (`--speculative-algorithm EAGLE` in SGLang, a `speculative-config` JSON block in vLLM) that point at the separate `mtp/` weights bundled in the model repo. This is where the recursive MTP layers come into play: at inference, a single recursive layer runs eight times to generate draft tokens, which the main model then verifies in a batch, speeding up generation without changing output quality.\n\nOnce the server is running, interaction is through the standard OpenAI-compatible chat completions API, using any base URL and a placeholder API key.\n\n## How does it integrate with Claude Code and Codex?\n\nIQuest-Q1 is built to slot into existing agentic coding harnesses rather than requiring a new client. For Claude Code, IQuest recommends version 2.1.140 and documents an `IQuest-Q1[1m]` model alias. Setting `ANTHROPIC_MODEL` and related environment variables to this alias redirects Claude Code’s requests through a gateway serving IQuest-Q1. The `[1m]` suffix is a labeling convenience only: the model card is explicit that it “does not change the 512K context limit,” so treat it as a context ceiling, not an upgrade.\n\nFor Codex CLI (version 0.142.0 recommended), the setup is more involved: it requires writing a custom `config.toml` and `model_catalog.json` into the Codex config directory, declaring a custom model provider that speaks the OpenAI Responses wire API, and setting the context window to 524,288 tokens with an auto-compact threshold around 419,430 tokens (80% of the window). Both integrations depend on a gateway in front of the SGLang or vLLM server that can translate between Anthropic Messages or OpenAI Responses formats and the model’s native API.\n\n## Is IQuest-Q1 worth running yourself?\n\nThat depends entirely on what hardware you already have access to and whether you need full control over an agentic coding model’s deployment. IQuest-Q1 competes on benchmarks the company frames against DeepSeek-V4-Flash/Pro releases, agentic coding suites like SWE-agent-style tasks, and an in-house CLI benchmark, though exact scores live in the model’s performance charts rather than as quoted numbers here. The architecture choices (MoE sparsity, hybrid attention, MTP speculative decoding) all point toward a team optimizing for serving efficiency at scale, not toward something designed for casual single-GPU use.\n\nFor teams that already run multi-GPU inference infrastructure and want an open-weight alternative to closed agentic coding models, the Docker images and documented launch commands make evaluation straightforward. For anyone without an 8-GPU node, this is one to watch for quantized community builds rather than something to deploy today.\n\n## Frequently Asked Questions\n\n### How much GPU memory does IQuest-Q1 need?\n\nIQuest doesn’t publish an exact VRAM figure, but the documented deployment uses tensor-parallel-8 across GPUs in bfloat16 for a 320B-parameter model, which implies a multi-GPU server with a combined memory pool well into the hundreds of gigabytes once KV cache and the MTP draft model are included.\n\n### Can I run IQuest-Q1 on a single GPU?\n\nNot with the officially documented configuration. Both the SGLang and vLLM commands specify `--tp-size 8` / `--tensor-parallel-size 8`, meaning the reference deployment assumes eight GPUs working together.\n\n### Does IQuest-Q1 support images or multimodal input?\n\nNo. The model card explicitly states this checkpoint is text-only, with no native image, audio, or video input capability.\n\n### What’s the difference between running it with and without MTP?\n\nMTP (multi-token prediction) enables speculative decoding: a small recursive draft layer predicts several tokens ahead, which the main model verifies in a batch. This speeds up generation throughput without changing the quality of outputs. Both SGLang and vLLM support launch configurations with and without it.\n\n### Can I use IQuest-Q1 with Claude Code or Codex instead of Anthropic or OpenAI models?\n\nYes, but only through a gateway that translates Anthropic Messages or OpenAI Responses API calls to the model’s native serving endpoint. IQuest documents environment variable and config file setups for both Claude Code (2.1.140) and Codex CLI (0.142.0).", "url": "https://wpnews.pro/news/iquest-q1-how-to-self-host-the-320b-agentic-coding-model", "canonical_source": "https://www.mindstudio.ai/blog/iquest-q1-local/", "published_at": "2026-10-03 00:00:00+00:00", "updated_at": "2026-10-03 18:38:40.564411+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-tools", "ai-infrastructure", "generative-ai"], "entities": ["IQuest", "IQuest-Q1", "SGLang", "vLLM", "Hugging Face", "Claude Code", "Codex CLI", "EAGLE"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/iquest-q1-how-to-self-host-the-320b-agentic-coding-model", "markdown": "https://wpnews.pro/news/iquest-q1-how-to-self-host-the-320b-agentic-coding-model.md", "text": "https://wpnews.pro/news/iquest-q1-how-to-self-host-the-320b-agentic-coding-model.txt", "jsonld": "https://wpnews.pro/news/iquest-q1-how-to-self-host-the-320b-agentic-coding-model.jsonld"}}