# How to Deploy IQuest-Q1 with SGLang or vLLM

> Source: <https://www.mindstudio.ai/blog/iquest-q1-local-deployment-sglang-vllm/>
> Published: 2026-09-30 00:00:00+00:00

# How to Deploy IQuest-Q1 with SGLang or vLLM

A practical guide to serving IQuest-Q1 locally with SGLang or vLLM, covering MTP speculative decoding, Docker setup, and tool-calling.

## What is IQuest-Q1 and what does it take to run it?

IQuest-Q1 is a 320-billion-parameter Mixture-of-Experts model from IQuest, built for agentic coding, reasoning, and multi-step tool use. Only about 15 billion parameters activate per token, thanks to a 256-expert routing setup that fires 8 experts at a time. Running it locally means multi-GPU tensor parallelism (the official examples use 8-way TP), a serving engine that understands its custom tool-call and reasoning parsers, and enough VRAM to hold an 88-layer transformer with a 512K-token context window. IQuest ships prebuilt Docker images for both SGLang and vLLM to make this tractable.

## TL;DR

- **IQuest-Q1** is a 320B-parameter MoE model with roughly 15B active parameters per token, an 88-layer transformer, 256 experts (8 active), and a**524,288-token context window** .
- The model uses a **hybrid attention pattern** (3 sliding-window attention layers per 1 full attention layer) plus partial RoPE, which keeps long-context inference more memory-efficient than dense full attention.
- IQuest provides ready-to-use **Docker images** for both SGLang (`sglang-iquest-q1:cu130` ) and vLLM (`vllm-iquest-q1:cu130` ), each configured for 8-way tensor parallelism.
- **Multi-Token Prediction (MTP)** is available as a speculative decoding draft model, implemented via EAGLE-style speculation in both engines, and can meaningfully speed up generation without changing output quality.
- Deploying with tool calling requires the model-specific `iquest_q1` parser flags for both**tool-call parsing** and**reasoning extraction** , not just generic OpenAI-compatible flags.
- The model integrates with **Claude Code** and**Codex CLI** through gateway configuration (Anthropic Messages API and OpenAI Responses API respectively), letting you swap it in as a drop-in agent backend.
- IQuest is explicit that this is an **early-stage release** : it’s text-only, can loop on failed attempts in real CLI tasks, and needs human review of generated code.

## Other agents ship a demo. Remy ships an app.

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

## What hardware and software do you need before starting?

The reference deployment commands from IQuest assume 8 GPUs with `--tp-size 8` (SGLang) or `--tensor-parallel-size 8` (vLLM). At 320B total parameters in bfloat16, the raw weights alone run into the hundreds of gigabytes, so this is squarely a multi-GPU, data-center-class deployment, not something you run on a single consumer card. You’ll also need:

- A CUDA 13.0-compatible environment (the Docker tags are explicitly `cu130` ).
- The Hugging Face CLI (`hf download` ) to pull`IQuestLab/IQuest-Q1` weights, including the`mtp` subdirectory if you plan to use speculative decoding.
- Docker, since IQuest publishes prebuilt images (`iquestlabworkspace/sglang-iquest-q1:cu130` and`iquestlabworkspace/vllm-iquest-q1:cu130` ) rather than requiring you to build the serving stack from source.
- `pip install openai` if you want to hit the server through the standard OpenAI Python client, since both engines expose an OpenAI-compatible chat completions endpoint.

## How do you deploy IQuest-Q1 with SGLang?

Pull the image, then launch the server pointed at the downloaded model root. The baseline command (no speculative decoding) looks like this:

```
docker pull iquestlabworkspace/sglang-iquest-q1:cu130

MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
python -u -m sglang.launch_server \
  --model-path "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tp-size 8 \
  --dtype bfloat16 \
  --attention-backend fa3 \
  --mem-fraction-static 0.85 \
  --disable-prefill-cuda-graph \
  --enable-metrics \
  --tool-call-parser iquest_q1 \
  --reasoning-parser iquest_q1 \
  --enable-torch-compile \
  --load-format fastsafetensors \
  --speculative-use-rejection-sampling
```

A few flags matter more than they look. `--attention-backend fa3` selects FlashAttention 3, which pairs with the model’s hybrid sliding-window/full-attention design. `--tool-call-parser iquest_q1` and `--reasoning-parser iquest_q1` are not optional if you want structured tool calls or reasoning traces extracted correctly, since IQuest-Q1 uses its own chat template format rather than a generic one. `--load-format fastsafetensors` speeds up weight loading for a model this size.

To add MTP-based speculative decoding, swap in the EAGLE configuration and point `--speculative-draft-model-path` at the `mtp` subfolder inside the downloaded model root:

```
python -u -m sglang.launch_server \
    --model-path "$MODEL_ROOT" \
    --served-model-name IQuest-Q1 \
    --tp-size 8 \
    --dtype bfloat16 \
    --attention-backend fa3 \
    --mem-fraction-static 0.85 \
    --disable-prefill-cuda-graph \
    --enable-metrics \
    --tool-call-parser iquest_q1 \
    --reasoning-parser iquest_q1 \
    --speculative-algorithm EAGLE \
    --speculative-num-steps 5 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 6 \
    --speculative-draft-model-path "$MODEL_ROOT/mtp" \
    --enable-torch-compile \
    --load-format fastsafetensors \
    --speculative-use-rejection-sampling \
    --speculative-draft-attention-backend fa3
```

## How do you deploy IQuest-Q1 with vLLM?

vLLM follows the same pattern with its own image and CLI syntax:

```
docker pull iquestlabworkspace/vllm-iquest-q1:cu130

MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
vllm serve "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tensor-parallel-size 8 \
  --reasoning-parser iquest_q1 \
  --enable-auto-tool-choice \
  --tool-call-parser iquest_q1
```

For speculative decoding, vLLM takes the EAGLE configuration as a JSON blob passed to `--speculative-config`, and it’s worth pairing that with `--enable-prefix-caching` given the long context window:

```
vllm serve "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tensor-parallel-size 8 \
  --reasoning-parser iquest_q1 \
  --enable-auto-tool-choice \
  --tool-call-parser iquest_q1 \
  --enable-prefix-caching \
  --speculative-config '{
    "method": "eagle",
    "model": "'"$MODEL_ROOT"'/mtp",
    "num_speculative_tokens": 5,
    "draft_sample_method": "probabilistic",
    "rejection_sample_method": "standard",
    "enforce_eager": false
  }'
```

Once either server is up, you talk to it through the standard OpenAI client:

``` python
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="sk-iquest")
response = client.chat.completions.create(
    model="IQuest-Q1",
    messages=[{"role": "user", "content": "Hello! Can you briefly introduce yourself?"}],
    temperature=1.0,
    top_p=0.95,
)
print(response.choices[0].message.content)
```

## Why does MTP speculative decoding matter here?

IQuest-Q1’s architecture includes dedicated Multi-Token Prediction layers: 2 independent MTP layers during training, collapsed into a single recursive layer applied 8 times at inference, with a 512-token MTP sliding window. In practice, this MTP module acts as the draft model for EAGLE-style speculative decoding. Instead of generating one token per forward pass through the full 320B-parameter model, the draft model proposes several tokens ahead, and the main model verifies them in a single batched pass with rejection sampling to preserve output correctness.

For a MoE model this large, that tradeoff is significant. Full forward passes through 88 layers with 8 active experts are expensive; verifying several speculative tokens at once is comparatively cheap. Both SGLang and vLLM expose this through near-identical parameters (`num_speculative_tokens` / `--speculative-num-draft-tokens`, both set to 5-6 in the reference configs), and both rely on rejection sampling to keep speculative output statistically equivalent to normal decoding, not just faster.

## How do you wire IQuest-Q1 into Claude Code or Codex?

Because IQuest-Q1 uses a custom tool-call format, plugging it into existing coding agents means routing through a gateway that translates Anthropic’s Messages API (for Claude Code) or OpenAI’s Responses API (for Codex) to your SGLang or vLLM endpoint.

For Claude Code, IQuest recommends version `2.1.140` and setting environment variables to point every internal model alias (Sonnet, Opus, Haiku, subagents) at `IQuest-Q1[1m]`, along with context and timeout settings tuned for the 512K window:

```
export ANTHROPIC_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_SONNET_MODEL="IQuest-Q1[1m]"
export CLAUDE_CODE_MAX_OUTPUT_TOKENS="131072"
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="524288"
export ANTHROPIC_BASE_URL="http://example-iquest-q1-link"
export ANTHROPIC_AUTH_TOKEN="sk-iquest-q1"
claude --model IQuest-Q1
```

For Codex CLI (version `0.142.0` recommended), IQuest provides a script that backs up and rewrites `config.toml` and `model_catalog.json` under `~/.codex`, setting `wire_api = "responses"`, a 524,288-token context window, and disabling sandbox approval prompts. This is meant for controlled local or internal environments, since it also disables sandboxing (`sandbox_mode = "danger-full-access"`), which is worth flagging before copying it into anything production-facing.

## Is IQuest-Q1 worth deploying locally?

That depends on what you’re optimizing for. It’s a serious agentic coding model with a 512K context window, tool-calling support, and reasoning traces built into its chat template, and it comes with first-party Docker images and speculative decoding support rather than a bare weights dump. That’s a meaningfully more complete deployment story than many open releases.

But it’s also, by IQuest’s own description, an early-stage model. The card explicitly warns that it’s text-only, that generated code needs review and testing, and that it can repeat failed attempts or overlook constraints on real-world CLI tasks. If you have the multi-GPU hardware to run a 320B-parameter MoE model, it’s a reasonable candidate to benchmark against other open agentic coding models. If you don’t have that hardware, or you need something turnkey and battle-tested, treat this as a model to watch rather than a drop-in replacement for your current stack.

## Frequently Asked Questions

### How much VRAM does IQuest-Q1 need?

## 
Plans first.
*Then code.*

Remy writes the spec, manages the build, and ships the app.

The model card doesn’t publish an exact VRAM figure, but the reference deployment commands use 8-way tensor parallelism across GPUs for a 320B-parameter model in bfloat16, which points to a multi-GPU, data-center-scale setup rather than anything running on a single card.

### Do I need the MTP weights to run IQuest-Q1?

No. The base deployment commands for both SGLang and vLLM work without them. MTP weights (found in the `mtp` subdirectory of the downloaded model) are only needed if you enable EAGLE-style speculative decoding for faster generation.

### Can I use IQuest-Q1 with images or other non-text input?

No. The model card lists text-only input as an explicit limitation. There is no native image, audio, or video capability in this checkpoint.

### What’s the difference between deploying with SGLang versus vLLM?

Functionally they achieve the same result: an OpenAI-compatible endpoint with tool calling and optional speculative decoding. The flags differ (SGLang uses discrete `--speculative-*` flags, vLLM takes a single JSON config), and IQuest provides separate Docker images for each, but both require the same `iquest_q1` tool-call and reasoning parsers.

### Is IQuest-Q1 production-ready?

IQuest describes it as an early-stage release with “substantial limitations in its capabilities and reliability.” It’s usable for experimentation and evaluation today, but the model card recommends human oversight on real-world CLI tasks and verification of any generated code before trusting it in production workflows.
