cd /news/large-language-models/how-to-deploy-iquest-q1-with-sglang-… · home › topics › large-language-models › article
[ARTICLE · art-142415] src=mindstudio.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

How to Deploy IQuest-Q1 with SGLang or vLLM

IQuest released IQuest-Q1, a 320-billion-parameter Mixture-of-Experts model with roughly 15 billion active parameters per token, an 88-layer transformer, 256 experts (8 active), and a 524,288-token context window, along with prebuilt Docker images for serving it on SGLang and vLLM. The deployment guide specifies 8-way tensor parallelism, CUDA 13.0, and model-specific iquest_q1 tool-call and reasoning parser flags, with Multi-Token Prediction available as an EAGLE-style speculative decoding draft model. IQuest labels the text-only release early-stage, noting the model can loop on failed attempts in real CLI tasks and requires human review of generated code.

by read8 min views1 publishedSep 30, 2026
How to Deploy IQuest-Q1 with SGLang or vLLM
Image: Mindstudio (auto-discovered)

A practical guide to serving IQuest-Q1 locally with SGLang or vLLM, covering MTP speculative decoding, Docker setup, and tool-calling.

What is IQuest-Q1 and what does it take to run it? #

IQuest-Q1 is a 320-billion-parameter Mixture-of-Experts model from IQuest, built for agentic coding, reasoning, and multi-step tool use. Only about 15 billion parameters activate per token, thanks to a 256-expert routing setup that fires 8 experts at a time. Running it locally means multi-GPU tensor parallelism (the official examples use 8-way TP), a serving engine that understands its custom tool-call and reasoning parsers, and enough VRAM to hold an 88-layer transformer with a 512K-token context window. IQuest ships prebuilt Docker images for both SGLang and vLLM to make this tractable.

TL;DR #

  • IQuest-Q1 is a 320B-parameter MoE model with roughly 15B active parameters per token, an 88-layer transformer, 256 experts (8 active), and a524,288-token context window .
  • The model uses a hybrid attention pattern (3 sliding-window attention layers per 1 full attention layer) plus partial RoPE, which keeps long-context inference more memory-efficient than dense full attention.
  • IQuest provides ready-to-use Docker images for both SGLang (sglang-iquest-q1:cu130 ) and vLLM (vllm-iquest-q1:cu130 ), each configured for 8-way tensor parallelism.
  • Multi-Token Prediction (MTP) is available as a speculative decoding draft model, implemented via EAGLE-style speculation in both engines, and can meaningfully speed up generation without changing output quality.
  • Deploying with tool calling requires the model-specific iquest_q1 parser flags for bothtool-call parsing andreasoning extraction , not just generic OpenAI-compatible flags.
  • The model integrates with Claude Code andCodex CLI through gateway configuration (Anthropic Messages API and OpenAI Responses API respectively), letting you swap it in as a drop-in agent backend.
  • IQuest is explicit that this is an early-stage release : it’s text-only, can loop on failed attempts in real CLI tasks, and needs human review of generated code.

Other agents ship a demo. Remy ships an app. #

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

What hardware and software do you need before starting? #

The reference deployment commands from IQuest assume 8 GPUs with --tp-size 8 (SGLang) or --tensor-parallel-size 8 (vLLM). At 320B total parameters in bfloat16, the raw weights alone run into the hundreds of gigabytes, so this is squarely a multi-GPU, data-center-class deployment, not something you run on a single consumer card. You’ll also need:

  • A CUDA 13.0-compatible environment (the Docker tags are explicitly cu130 ).
  • The Hugging Face CLI (hf download ) to pullIQuestLab/IQuest-Q1 weights, including themtp subdirectory if you plan to use speculative decoding.
  • Docker, since IQuest publishes prebuilt images (iquestlabworkspace/sglang-iquest-q1:cu130 andiquestlabworkspace/vllm-iquest-q1:cu130 ) rather than requiring you to build the serving stack from source.
  • pip install openai if you want to hit the server through the standard OpenAI Python client, since both engines expose an OpenAI-compatible chat completions endpoint.

How do you deploy IQuest-Q1 with SGLang? #

Pull the image, then launch the server pointed at the downloaded model root. The baseline command (no speculative decoding) looks like this:

docker pull iquestlabworkspace/sglang-iquest-q1:cu130

MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
python -u -m sglang.launch_server \
  --model-path "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tp-size 8 \
  --dtype bfloat16 \
  --attention-backend fa3 \
  --mem-fraction-static 0.85 \
  --disable-prefill-cuda-graph \
  --enable-metrics \
  --tool-call-parser iquest_q1 \
  --reasoning-parser iquest_q1 \
  --enable-torch-compile \
  --load-format fastsafetensors \
  --speculative-use-rejection-sampling

A few flags matter more than they look. --attention-backend fa3 selects FlashAttention 3, which pairs with the model’s hybrid sliding-window/full-attention design. --tool-call-parser iquest_q1 and --reasoning-parser iquest_q1 are not optional if you want structured tool calls or reasoning traces extracted correctly, since IQuest-Q1 uses its own chat template format rather than a generic one. --load-format fastsafetensors speeds up weight for a model this size.

To add MTP-based speculative decoding, swap in the EAGLE configuration and point --speculative-draft-model-path at the mtp subfolder inside the downloaded model root:

python -u -m sglang.launch_server \
    --model-path "$MODEL_ROOT" \
    --served-model-name IQuest-Q1 \
    --tp-size 8 \
    --dtype bfloat16 \
    --attention-backend fa3 \
    --mem-fraction-static 0.85 \
    --disable-prefill-cuda-graph \
    --enable-metrics \
    --tool-call-parser iquest_q1 \
    --reasoning-parser iquest_q1 \
    --speculative-algorithm EAGLE \
    --speculative-num-steps 5 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 6 \
    --speculative-draft-model-path "$MODEL_ROOT/mtp" \
    --enable-torch-compile \
    --load-format fastsafetensors \
    --speculative-use-rejection-sampling \
    --speculative-draft-attention-backend fa3

How do you deploy IQuest-Q1 with vLLM? #

vLLM follows the same pattern with its own image and CLI syntax:

docker pull iquestlabworkspace/vllm-iquest-q1:cu130

MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
vllm serve "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tensor-parallel-size 8 \
  --reasoning-parser iquest_q1 \
  --enable-auto-tool-choice \
  --tool-call-parser iquest_q1

For speculative decoding, vLLM takes the EAGLE configuration as a JSON blob passed to --speculative-config, and it’s worth pairing that with --enable-prefix-caching given the long context window:

vllm serve "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tensor-parallel-size 8 \
  --reasoning-parser iquest_q1 \
  --enable-auto-tool-choice \
  --tool-call-parser iquest_q1 \
  --enable-prefix-caching \
  --speculative-config '{
    "method": "eagle",
    "model": "'"$MODEL_ROOT"'/mtp",
    "num_speculative_tokens": 5,
    "draft_sample_method": "probabilistic",
    "rejection_sample_method": "standard",
    "enforce_eager": false
  }'

Once either server is up, you talk to it through the standard OpenAI client:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="sk-iquest")
response = client.chat.completions.create(
    model="IQuest-Q1",
    messages=[{"role": "user", "content": "Hello! Can you briefly introduce yourself?"}],
    temperature=1.0,
    top_p=0.95,
)
print(response.choices[0].message.content)

Why does MTP speculative decoding matter here? #

IQuest-Q1’s architecture includes dedicated Multi-Token Prediction layers: 2 independent MTP layers during training, collapsed into a single recursive layer applied 8 times at inference, with a 512-token MTP sliding window. In practice, this MTP module acts as the draft model for EAGLE-style speculative decoding. Instead of generating one token per forward pass through the full 320B-parameter model, the draft model proposes several tokens ahead, and the main model verifies them in a single batched pass with rejection sampling to preserve output correctness.

For a MoE model this large, that tradeoff is significant. Full forward passes through 88 layers with 8 active experts are expensive; verifying several speculative tokens at once is comparatively cheap. Both SGLang and vLLM expose this through near-identical parameters (num_speculative_tokens / --speculative-num-draft-tokens, both set to 5-6 in the reference configs), and both rely on rejection sampling to keep speculative output statistically equivalent to normal decoding, not just faster.

How do you wire IQuest-Q1 into Claude Code or Codex? #

Because IQuest-Q1 uses a custom tool-call format, plugging it into existing coding agents means routing through a gateway that translates Anthropic’s Messages API (for Claude Code) or OpenAI’s Responses API (for Codex) to your SGLang or vLLM endpoint.

For Claude Code, IQuest recommends version 2.1.140 and setting environment variables to point every internal model alias (Sonnet, Opus, Haiku, subagents) at IQuest-Q1[1m], along with context and timeout settings tuned for the 512K window:

export ANTHROPIC_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_SONNET_MODEL="IQuest-Q1[1m]"
export CLAUDE_CODE_MAX_OUTPUT_TOKENS="131072"
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="524288"
export ANTHROPIC_BASE_URL="http://example-iquest-q1-link"
export ANTHROPIC_AUTH_TOKEN="sk-iquest-q1"
claude --model IQuest-Q1

For Codex CLI (version 0.142.0 recommended), IQuest provides a script that backs up and rewrites config.toml and model_catalog.json under ~/.codex, setting wire_api = "responses", a 524,288-token context window, and disabling sandbox approval prompts. This is meant for controlled local or internal environments, since it also disables sandboxing (sandbox_mode = "danger-full-access"), which is worth flagging before copying it into anything production-facing.

Is IQuest-Q1 worth deploying locally? #

That depends on what you’re optimizing for. It’s a serious agentic coding model with a 512K context window, tool-calling support, and reasoning traces built into its chat template, and it comes with first-party Docker images and speculative decoding support rather than a bare weights dump. That’s a meaningfully more complete deployment story than many open releases.

But it’s also, by IQuest’s own description, an early-stage model. The card explicitly warns that it’s text-only, that generated code needs review and testing, and that it can repeat failed attempts or overlook constraints on real-world CLI tasks. If you have the multi-GPU hardware to run a 320B-parameter MoE model, it’s a reasonable candidate to benchmark against other open agentic coding models. If you don’t have that hardware, or you need something turnkey and battle-tested, treat this as a model to watch rather than a drop-in replacement for your current stack.

Frequently Asked Questions #

How much VRAM does IQuest-Q1 need?

#

Plans first. Then code.

Remy writes the spec, manages the build, and ships the app.

The model card doesn’t publish an exact VRAM figure, but the reference deployment commands use 8-way tensor parallelism across GPUs for a 320B-parameter model in bfloat16, which points to a multi-GPU, data-center-scale setup rather than anything running on a single card.

Do I need the MTP weights to run IQuest-Q1?

No. The base deployment commands for both SGLang and vLLM work without them. MTP weights (found in the mtp subdirectory of the downloaded model) are only needed if you enable EAGLE-style speculative decoding for faster generation.

Can I use IQuest-Q1 with images or other non-text input?

No. The model card lists text-only input as an explicit limitation. There is no native image, audio, or video capability in this checkpoint.

What’s the difference between deploying with SGLang versus vLLM?

Functionally they achieve the same result: an OpenAI-compatible endpoint with tool calling and optional speculative decoding. The flags differ (SGLang uses discrete --speculative-* flags, vLLM takes a single JSON config), and IQuest provides separate Docker images for each, but both require the same iquest_q1 tool-call and reasoning parsers.

Is IQuest-Q1 production-ready?

IQuest describes it as an early-stage release with “substantial limitations in its capabilities and reliability.” It’s usable for experimentation and evaluation today, but the model card recommends human oversight on real-world CLI tasks and verification of any generated code before trusting it in production workflows.

── more in #large-language-models 4 stories · sorted by recency
── more on @iquest 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-deploy-iquest…] indexed:0 read:8min 2026-09-30 · —