How to Deploy IQuest-Q1 with SGLang or vLLM IQuest released IQuest-Q1, a 320-billion-parameter Mixture-of-Experts model with roughly 15 billion active parameters per token, an 88-layer transformer, 256 experts (8 active), and a 524,288-token context window, along with prebuilt Docker images for serving it on SGLang and vLLM. The deployment guide specifies 8-way tensor parallelism, CUDA 13.0, and model-specific iquest_q1 tool-call and reasoning parser flags, with Multi-Token Prediction available as an EAGLE-style speculative decoding draft model. IQuest labels the text-only release early-stage, noting the model can loop on failed attempts in real CLI tasks and requires human review of generated code. How to Deploy IQuest-Q1 with SGLang or vLLM A practical guide to serving IQuest-Q1 locally with SGLang or vLLM, covering MTP speculative decoding, Docker setup, and tool-calling. What is IQuest-Q1 and what does it take to run it? IQuest-Q1 is a 320-billion-parameter Mixture-of-Experts model from IQuest, built for agentic coding, reasoning, and multi-step tool use. Only about 15 billion parameters activate per token, thanks to a 256-expert routing setup that fires 8 experts at a time. Running it locally means multi-GPU tensor parallelism the official examples use 8-way TP , a serving engine that understands its custom tool-call and reasoning parsers, and enough VRAM to hold an 88-layer transformer with a 512K-token context window. IQuest ships prebuilt Docker images for both SGLang and vLLM to make this tractable. TL;DR - IQuest-Q1 is a 320B-parameter MoE model with roughly 15B active parameters per token, an 88-layer transformer, 256 experts 8 active , and a 524,288-token context window . - The model uses a hybrid attention pattern 3 sliding-window attention layers per 1 full attention layer plus partial RoPE, which keeps long-context inference more memory-efficient than dense full attention. - IQuest provides ready-to-use Docker images for both SGLang sglang-iquest-q1:cu130 and vLLM vllm-iquest-q1:cu130 , each configured for 8-way tensor parallelism. - Multi-Token Prediction MTP is available as a speculative decoding draft model, implemented via EAGLE-style speculation in both engines, and can meaningfully speed up generation without changing output quality. - Deploying with tool calling requires the model-specific iquest q1 parser flags for both tool-call parsing and reasoning extraction , not just generic OpenAI-compatible flags. - The model integrates with Claude Code and Codex CLI through gateway configuration Anthropic Messages API and OpenAI Responses API respectively , letting you swap it in as a drop-in agent backend. - IQuest is explicit that this is an early-stage release : it’s text-only, can loop on failed attempts in real CLI tasks, and needs human review of generated code. Other agents ship a demo. Remy ships an app. Real backend. Real database. Real auth. Real plumbing. Remy has it all. What hardware and software do you need before starting? The reference deployment commands from IQuest assume 8 GPUs with --tp-size 8 SGLang or --tensor-parallel-size 8 vLLM . At 320B total parameters in bfloat16, the raw weights alone run into the hundreds of gigabytes, so this is squarely a multi-GPU, data-center-class deployment, not something you run on a single consumer card. You’ll also need: - A CUDA 13.0-compatible environment the Docker tags are explicitly cu130 . - The Hugging Face CLI hf download to pull IQuestLab/IQuest-Q1 weights, including the mtp subdirectory if you plan to use speculative decoding. - Docker, since IQuest publishes prebuilt images iquestlabworkspace/sglang-iquest-q1:cu130 and iquestlabworkspace/vllm-iquest-q1:cu130 rather than requiring you to build the serving stack from source. - pip install openai if you want to hit the server through the standard OpenAI Python client, since both engines expose an OpenAI-compatible chat completions endpoint. How do you deploy IQuest-Q1 with SGLang? Pull the image, then launch the server pointed at the downloaded model root. The baseline command no speculative decoding looks like this: docker pull iquestlabworkspace/sglang-iquest-q1:cu130 MODEL ROOT="$ hf download IQuestLab/IQuest-Q1 --quiet " && \ python -u -m sglang.launch server \ --model-path "$MODEL ROOT" \ --served-model-name IQuest-Q1 \ --tp-size 8 \ --dtype bfloat16 \ --attention-backend fa3 \ --mem-fraction-static 0.85 \ --disable-prefill-cuda-graph \ --enable-metrics \ --tool-call-parser iquest q1 \ --reasoning-parser iquest q1 \ --enable-torch-compile \ --load-format fastsafetensors \ --speculative-use-rejection-sampling A few flags matter more than they look. --attention-backend fa3 selects FlashAttention 3, which pairs with the model’s hybrid sliding-window/full-attention design. --tool-call-parser iquest q1 and --reasoning-parser iquest q1 are not optional if you want structured tool calls or reasoning traces extracted correctly, since IQuest-Q1 uses its own chat template format rather than a generic one. --load-format fastsafetensors speeds up weight loading for a model this size. To add MTP-based speculative decoding, swap in the EAGLE configuration and point --speculative-draft-model-path at the mtp subfolder inside the downloaded model root: python -u -m sglang.launch server \ --model-path "$MODEL ROOT" \ --served-model-name IQuest-Q1 \ --tp-size 8 \ --dtype bfloat16 \ --attention-backend fa3 \ --mem-fraction-static 0.85 \ --disable-prefill-cuda-graph \ --enable-metrics \ --tool-call-parser iquest q1 \ --reasoning-parser iquest q1 \ --speculative-algorithm EAGLE \ --speculative-num-steps 5 \ --speculative-eagle-topk 1 \ --speculative-num-draft-tokens 6 \ --speculative-draft-model-path "$MODEL ROOT/mtp" \ --enable-torch-compile \ --load-format fastsafetensors \ --speculative-use-rejection-sampling \ --speculative-draft-attention-backend fa3 How do you deploy IQuest-Q1 with vLLM? vLLM follows the same pattern with its own image and CLI syntax: docker pull iquestlabworkspace/vllm-iquest-q1:cu130 MODEL ROOT="$ hf download IQuestLab/IQuest-Q1 --quiet " && \ vllm serve "$MODEL ROOT" \ --served-model-name IQuest-Q1 \ --tensor-parallel-size 8 \ --reasoning-parser iquest q1 \ --enable-auto-tool-choice \ --tool-call-parser iquest q1 For speculative decoding, vLLM takes the EAGLE configuration as a JSON blob passed to --speculative-config , and it’s worth pairing that with --enable-prefix-caching given the long context window: vllm serve "$MODEL ROOT" \ --served-model-name IQuest-Q1 \ --tensor-parallel-size 8 \ --reasoning-parser iquest q1 \ --enable-auto-tool-choice \ --tool-call-parser iquest q1 \ --enable-prefix-caching \ --speculative-config '{ "method": "eagle", "model": "'"$MODEL ROOT"'/mtp", "num speculative tokens": 5, "draft sample method": "probabilistic", "rejection sample method": "standard", "enforce eager": false }' Once either server is up, you talk to it through the standard OpenAI client: python from openai import OpenAI client = OpenAI base url="http://127.0.0.1:8000/v1", api key="sk-iquest" response = client.chat.completions.create model="IQuest-Q1", messages= {"role": "user", "content": "Hello Can you briefly introduce yourself?"} , temperature=1.0, top p=0.95, print response.choices 0 .message.content Why does MTP speculative decoding matter here? IQuest-Q1’s architecture includes dedicated Multi-Token Prediction layers: 2 independent MTP layers during training, collapsed into a single recursive layer applied 8 times at inference, with a 512-token MTP sliding window. In practice, this MTP module acts as the draft model for EAGLE-style speculative decoding. Instead of generating one token per forward pass through the full 320B-parameter model, the draft model proposes several tokens ahead, and the main model verifies them in a single batched pass with rejection sampling to preserve output correctness. For a MoE model this large, that tradeoff is significant. Full forward passes through 88 layers with 8 active experts are expensive; verifying several speculative tokens at once is comparatively cheap. Both SGLang and vLLM expose this through near-identical parameters num speculative tokens / --speculative-num-draft-tokens , both set to 5-6 in the reference configs , and both rely on rejection sampling to keep speculative output statistically equivalent to normal decoding, not just faster. How do you wire IQuest-Q1 into Claude Code or Codex? Because IQuest-Q1 uses a custom tool-call format, plugging it into existing coding agents means routing through a gateway that translates Anthropic’s Messages API for Claude Code or OpenAI’s Responses API for Codex to your SGLang or vLLM endpoint. For Claude Code, IQuest recommends version 2.1.140 and setting environment variables to point every internal model alias Sonnet, Opus, Haiku, subagents at IQuest-Q1 1m , along with context and timeout settings tuned for the 512K window: export ANTHROPIC MODEL="IQuest-Q1 1m " export ANTHROPIC DEFAULT SONNET MODEL="IQuest-Q1 1m " export CLAUDE CODE MAX OUTPUT TOKENS="131072" export CLAUDE CODE AUTO COMPACT WINDOW="524288" export ANTHROPIC BASE URL="http://example-iquest-q1-link" export ANTHROPIC AUTH TOKEN="sk-iquest-q1" claude --model IQuest-Q1 For Codex CLI version 0.142.0 recommended , IQuest provides a script that backs up and rewrites config.toml and model catalog.json under ~/.codex , setting wire api = "responses" , a 524,288-token context window, and disabling sandbox approval prompts. This is meant for controlled local or internal environments, since it also disables sandboxing sandbox mode = "danger-full-access" , which is worth flagging before copying it into anything production-facing. Is IQuest-Q1 worth deploying locally? That depends on what you’re optimizing for. It’s a serious agentic coding model with a 512K context window, tool-calling support, and reasoning traces built into its chat template, and it comes with first-party Docker images and speculative decoding support rather than a bare weights dump. That’s a meaningfully more complete deployment story than many open releases. But it’s also, by IQuest’s own description, an early-stage model. The card explicitly warns that it’s text-only, that generated code needs review and testing, and that it can repeat failed attempts or overlook constraints on real-world CLI tasks. If you have the multi-GPU hardware to run a 320B-parameter MoE model, it’s a reasonable candidate to benchmark against other open agentic coding models. If you don’t have that hardware, or you need something turnkey and battle-tested, treat this as a model to watch rather than a drop-in replacement for your current stack. Frequently Asked Questions How much VRAM does IQuest-Q1 need? Plans first. Then code. Remy writes the spec, manages the build, and ships the app. The model card doesn’t publish an exact VRAM figure, but the reference deployment commands use 8-way tensor parallelism across GPUs for a 320B-parameter model in bfloat16, which points to a multi-GPU, data-center-scale setup rather than anything running on a single card. Do I need the MTP weights to run IQuest-Q1? No. The base deployment commands for both SGLang and vLLM work without them. MTP weights found in the mtp subdirectory of the downloaded model are only needed if you enable EAGLE-style speculative decoding for faster generation. Can I use IQuest-Q1 with images or other non-text input? No. The model card lists text-only input as an explicit limitation. There is no native image, audio, or video capability in this checkpoint. What’s the difference between deploying with SGLang versus vLLM? Functionally they achieve the same result: an OpenAI-compatible endpoint with tool calling and optional speculative decoding. The flags differ SGLang uses discrete --speculative- flags, vLLM takes a single JSON config , and IQuest provides separate Docker images for each, but both require the same iquest q1 tool-call and reasoning parsers. Is IQuest-Q1 production-ready? IQuest describes it as an early-stage release with “substantial limitations in its capabilities and reliability.” It’s usable for experimentation and evaluation today, but the model card recommends human oversight on real-world CLI tasks and verification of any generated code before trusting it in production workflows.