{"slug": "how-to-deploy-iquest-q1-with-sglang-or-vllm", "title": "How to Deploy IQuest-Q1 with SGLang or vLLM", "summary": "IQuest released IQuest-Q1, a 320-billion-parameter Mixture-of-Experts model with roughly 15 billion active parameters per token, an 88-layer transformer, 256 experts (8 active), and a 524,288-token context window, along with prebuilt Docker images for serving it on SGLang and vLLM. The deployment guide specifies 8-way tensor parallelism, CUDA 13.0, and model-specific iquest_q1 tool-call and reasoning parser flags, with Multi-Token Prediction available as an EAGLE-style speculative decoding draft model. IQuest labels the text-only release early-stage, noting the model can loop on failed attempts in real CLI tasks and requires human review of generated code.", "body_md": "# How to Deploy IQuest-Q1 with SGLang or vLLM\n\nA practical guide to serving IQuest-Q1 locally with SGLang or vLLM, covering MTP speculative decoding, Docker setup, and tool-calling.\n\n## What is IQuest-Q1 and what does it take to run it?\n\nIQuest-Q1 is a 320-billion-parameter Mixture-of-Experts model from IQuest, built for agentic coding, reasoning, and multi-step tool use. Only about 15 billion parameters activate per token, thanks to a 256-expert routing setup that fires 8 experts at a time. Running it locally means multi-GPU tensor parallelism (the official examples use 8-way TP), a serving engine that understands its custom tool-call and reasoning parsers, and enough VRAM to hold an 88-layer transformer with a 512K-token context window. IQuest ships prebuilt Docker images for both SGLang and vLLM to make this tractable.\n\n## TL;DR\n\n- **IQuest-Q1** is a 320B-parameter MoE model with roughly 15B active parameters per token, an 88-layer transformer, 256 experts (8 active), and a**524,288-token context window** .\n- The model uses a **hybrid attention pattern** (3 sliding-window attention layers per 1 full attention layer) plus partial RoPE, which keeps long-context inference more memory-efficient than dense full attention.\n- IQuest provides ready-to-use **Docker images** for both SGLang (`sglang-iquest-q1:cu130` ) and vLLM (`vllm-iquest-q1:cu130` ), each configured for 8-way tensor parallelism.\n- **Multi-Token Prediction (MTP)** is available as a speculative decoding draft model, implemented via EAGLE-style speculation in both engines, and can meaningfully speed up generation without changing output quality.\n- Deploying with tool calling requires the model-specific `iquest_q1` parser flags for both**tool-call parsing** and**reasoning extraction** , not just generic OpenAI-compatible flags.\n- The model integrates with **Claude Code** and**Codex CLI** through gateway configuration (Anthropic Messages API and OpenAI Responses API respectively), letting you swap it in as a drop-in agent backend.\n- IQuest is explicit that this is an **early-stage release** : it’s text-only, can loop on failed attempts in real CLI tasks, and needs human review of generated code.\n\n## Other agents ship a demo. Remy ships an app.\n\nReal backend. Real database. Real auth. Real plumbing. Remy has it all.\n\n## What hardware and software do you need before starting?\n\nThe reference deployment commands from IQuest assume 8 GPUs with `--tp-size 8` (SGLang) or `--tensor-parallel-size 8` (vLLM). At 320B total parameters in bfloat16, the raw weights alone run into the hundreds of gigabytes, so this is squarely a multi-GPU, data-center-class deployment, not something you run on a single consumer card. You’ll also need:\n\n- A CUDA 13.0-compatible environment (the Docker tags are explicitly `cu130` ).\n- The Hugging Face CLI (`hf download` ) to pull`IQuestLab/IQuest-Q1` weights, including the`mtp` subdirectory if you plan to use speculative decoding.\n- Docker, since IQuest publishes prebuilt images (`iquestlabworkspace/sglang-iquest-q1:cu130` and`iquestlabworkspace/vllm-iquest-q1:cu130` ) rather than requiring you to build the serving stack from source.\n- `pip install openai` if you want to hit the server through the standard OpenAI Python client, since both engines expose an OpenAI-compatible chat completions endpoint.\n\n## How do you deploy IQuest-Q1 with SGLang?\n\nPull the image, then launch the server pointed at the downloaded model root. The baseline command (no speculative decoding) looks like this:\n\n```\ndocker pull iquestlabworkspace/sglang-iquest-q1:cu130\n\nMODEL_ROOT=\"$(hf download IQuestLab/IQuest-Q1 --quiet)\" && \\\npython -u -m sglang.launch_server \\\n  --model-path \"$MODEL_ROOT\" \\\n  --served-model-name IQuest-Q1 \\\n  --tp-size 8 \\\n  --dtype bfloat16 \\\n  --attention-backend fa3 \\\n  --mem-fraction-static 0.85 \\\n  --disable-prefill-cuda-graph \\\n  --enable-metrics \\\n  --tool-call-parser iquest_q1 \\\n  --reasoning-parser iquest_q1 \\\n  --enable-torch-compile \\\n  --load-format fastsafetensors \\\n  --speculative-use-rejection-sampling\n```\n\nA few flags matter more than they look. `--attention-backend fa3` selects FlashAttention 3, which pairs with the model’s hybrid sliding-window/full-attention design. `--tool-call-parser iquest_q1` and `--reasoning-parser iquest_q1` are not optional if you want structured tool calls or reasoning traces extracted correctly, since IQuest-Q1 uses its own chat template format rather than a generic one. `--load-format fastsafetensors` speeds up weight loading for a model this size.\n\nTo add MTP-based speculative decoding, swap in the EAGLE configuration and point `--speculative-draft-model-path` at the `mtp` subfolder inside the downloaded model root:\n\n```\npython -u -m sglang.launch_server \\\n    --model-path \"$MODEL_ROOT\" \\\n    --served-model-name IQuest-Q1 \\\n    --tp-size 8 \\\n    --dtype bfloat16 \\\n    --attention-backend fa3 \\\n    --mem-fraction-static 0.85 \\\n    --disable-prefill-cuda-graph \\\n    --enable-metrics \\\n    --tool-call-parser iquest_q1 \\\n    --reasoning-parser iquest_q1 \\\n    --speculative-algorithm EAGLE \\\n    --speculative-num-steps 5 \\\n    --speculative-eagle-topk 1 \\\n    --speculative-num-draft-tokens 6 \\\n    --speculative-draft-model-path \"$MODEL_ROOT/mtp\" \\\n    --enable-torch-compile \\\n    --load-format fastsafetensors \\\n    --speculative-use-rejection-sampling \\\n    --speculative-draft-attention-backend fa3\n```\n\n## How do you deploy IQuest-Q1 with vLLM?\n\nvLLM follows the same pattern with its own image and CLI syntax:\n\n```\ndocker pull iquestlabworkspace/vllm-iquest-q1:cu130\n\nMODEL_ROOT=\"$(hf download IQuestLab/IQuest-Q1 --quiet)\" && \\\nvllm serve \"$MODEL_ROOT\" \\\n  --served-model-name IQuest-Q1 \\\n  --tensor-parallel-size 8 \\\n  --reasoning-parser iquest_q1 \\\n  --enable-auto-tool-choice \\\n  --tool-call-parser iquest_q1\n```\n\nFor speculative decoding, vLLM takes the EAGLE configuration as a JSON blob passed to `--speculative-config`, and it’s worth pairing that with `--enable-prefix-caching` given the long context window:\n\n```\nvllm serve \"$MODEL_ROOT\" \\\n  --served-model-name IQuest-Q1 \\\n  --tensor-parallel-size 8 \\\n  --reasoning-parser iquest_q1 \\\n  --enable-auto-tool-choice \\\n  --tool-call-parser iquest_q1 \\\n  --enable-prefix-caching \\\n  --speculative-config '{\n    \"method\": \"eagle\",\n    \"model\": \"'\"$MODEL_ROOT\"'/mtp\",\n    \"num_speculative_tokens\": 5,\n    \"draft_sample_method\": \"probabilistic\",\n    \"rejection_sample_method\": \"standard\",\n    \"enforce_eager\": false\n  }'\n```\n\nOnce either server is up, you talk to it through the standard OpenAI client:\n\n``` python\nfrom openai import OpenAI\n\nclient = OpenAI(base_url=\"http://127.0.0.1:8000/v1\", api_key=\"sk-iquest\")\nresponse = client.chat.completions.create(\n    model=\"IQuest-Q1\",\n    messages=[{\"role\": \"user\", \"content\": \"Hello! Can you briefly introduce yourself?\"}],\n    temperature=1.0,\n    top_p=0.95,\n)\nprint(response.choices[0].message.content)\n```\n\n## Why does MTP speculative decoding matter here?\n\nIQuest-Q1’s architecture includes dedicated Multi-Token Prediction layers: 2 independent MTP layers during training, collapsed into a single recursive layer applied 8 times at inference, with a 512-token MTP sliding window. In practice, this MTP module acts as the draft model for EAGLE-style speculative decoding. Instead of generating one token per forward pass through the full 320B-parameter model, the draft model proposes several tokens ahead, and the main model verifies them in a single batched pass with rejection sampling to preserve output correctness.\n\nFor a MoE model this large, that tradeoff is significant. Full forward passes through 88 layers with 8 active experts are expensive; verifying several speculative tokens at once is comparatively cheap. Both SGLang and vLLM expose this through near-identical parameters (`num_speculative_tokens` / `--speculative-num-draft-tokens`, both set to 5-6 in the reference configs), and both rely on rejection sampling to keep speculative output statistically equivalent to normal decoding, not just faster.\n\n## How do you wire IQuest-Q1 into Claude Code or Codex?\n\nBecause IQuest-Q1 uses a custom tool-call format, plugging it into existing coding agents means routing through a gateway that translates Anthropic’s Messages API (for Claude Code) or OpenAI’s Responses API (for Codex) to your SGLang or vLLM endpoint.\n\nFor Claude Code, IQuest recommends version `2.1.140` and setting environment variables to point every internal model alias (Sonnet, Opus, Haiku, subagents) at `IQuest-Q1[1m]`, along with context and timeout settings tuned for the 512K window:\n\n```\nexport ANTHROPIC_MODEL=\"IQuest-Q1[1m]\"\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"IQuest-Q1[1m]\"\nexport CLAUDE_CODE_MAX_OUTPUT_TOKENS=\"131072\"\nexport CLAUDE_CODE_AUTO_COMPACT_WINDOW=\"524288\"\nexport ANTHROPIC_BASE_URL=\"http://example-iquest-q1-link\"\nexport ANTHROPIC_AUTH_TOKEN=\"sk-iquest-q1\"\nclaude --model IQuest-Q1\n```\n\nFor Codex CLI (version `0.142.0` recommended), IQuest provides a script that backs up and rewrites `config.toml` and `model_catalog.json` under `~/.codex`, setting `wire_api = \"responses\"`, a 524,288-token context window, and disabling sandbox approval prompts. This is meant for controlled local or internal environments, since it also disables sandboxing (`sandbox_mode = \"danger-full-access\"`), which is worth flagging before copying it into anything production-facing.\n\n## Is IQuest-Q1 worth deploying locally?\n\nThat depends on what you’re optimizing for. It’s a serious agentic coding model with a 512K context window, tool-calling support, and reasoning traces built into its chat template, and it comes with first-party Docker images and speculative decoding support rather than a bare weights dump. That’s a meaningfully more complete deployment story than many open releases.\n\nBut it’s also, by IQuest’s own description, an early-stage model. The card explicitly warns that it’s text-only, that generated code needs review and testing, and that it can repeat failed attempts or overlook constraints on real-world CLI tasks. If you have the multi-GPU hardware to run a 320B-parameter MoE model, it’s a reasonable candidate to benchmark against other open agentic coding models. If you don’t have that hardware, or you need something turnkey and battle-tested, treat this as a model to watch rather than a drop-in replacement for your current stack.\n\n## Frequently Asked Questions\n\n### How much VRAM does IQuest-Q1 need?\n\n## \nPlans first.\n*Then code.*\n\nRemy writes the spec, manages the build, and ships the app.\n\nThe model card doesn’t publish an exact VRAM figure, but the reference deployment commands use 8-way tensor parallelism across GPUs for a 320B-parameter model in bfloat16, which points to a multi-GPU, data-center-scale setup rather than anything running on a single card.\n\n### Do I need the MTP weights to run IQuest-Q1?\n\nNo. The base deployment commands for both SGLang and vLLM work without them. MTP weights (found in the `mtp` subdirectory of the downloaded model) are only needed if you enable EAGLE-style speculative decoding for faster generation.\n\n### Can I use IQuest-Q1 with images or other non-text input?\n\nNo. The model card lists text-only input as an explicit limitation. There is no native image, audio, or video capability in this checkpoint.\n\n### What’s the difference between deploying with SGLang versus vLLM?\n\nFunctionally they achieve the same result: an OpenAI-compatible endpoint with tool calling and optional speculative decoding. The flags differ (SGLang uses discrete `--speculative-*` flags, vLLM takes a single JSON config), and IQuest provides separate Docker images for each, but both require the same `iquest_q1` tool-call and reasoning parsers.\n\n### Is IQuest-Q1 production-ready?\n\nIQuest describes it as an early-stage release with “substantial limitations in its capabilities and reliability.” It’s usable for experimentation and evaluation today, but the model card recommends human oversight on real-world CLI tasks and verification of any generated code before trusting it in production workflows.", "url": "https://wpnews.pro/news/how-to-deploy-iquest-q1-with-sglang-or-vllm", "canonical_source": "https://www.mindstudio.ai/blog/iquest-q1-local-deployment-sglang-vllm/", "published_at": "2026-09-30 00:00:00+00:00", "updated_at": "2026-09-30 10:18:27.583511+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-agents", "ai-tools", "mlops"], "entities": ["IQuest", "IQuest-Q1", "SGLang", "vLLM", "Hugging Face", "Claude Code", "Codex CLI", "Docker"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-deploy-iquest-q1-with-sglang-or-vllm", "markdown": "https://wpnews.pro/news/how-to-deploy-iquest-q1-with-sglang-or-vllm.md", "text": "https://wpnews.pro/news/how-to-deploy-iquest-q1-with-sglang-or-vllm.txt", "jsonld": "https://wpnews.pro/news/how-to-deploy-iquest-q1-with-sglang-or-vllm.jsonld"}}