LLM Scaler โ€“ LLM Support for Intel's Arc Pro B60 and B70 GPUs Intel has released LLM Scaler, a GenAI solution for text, image, and video generation optimized for Intel Arc Pro B60 and B70 GPUs, with the latest version intel/llm-scaler-vllm:0.21.0-b2 adding Multi-token Prediction and Lora Serving for Qwen3.6-27B, Qwen3.6-35B-A3B, gemma-4-31B-it, and gemma-4-26B-A4B-it models, plus FP8 per-block quantization support. The solution leverages frameworks such as vLLM, ComfyUI, SGLang Diffusion, and Xinference to deliver performance for state-of-the-art GenAI models. LLM Scaler is an GenAI solution for text generation, image generation, video generation etc. running on Intelยฎ Arcโ„ข Pro B60 and B70 GPUs. LLM Scalar leverages standard frameworks such as vLLM, ComfyUI, SGLang Diffusion, Xinference etc and ensures the best performance for State-of-Art GenAI models running on Arc Pro B60/B70 GPUs. - ๐Ÿ”ฅ 2026.08 We released intel/llm-scaler-vllm:0.21.0-b2 to support Multi-token Prediction MTP and Lora Serving for Qwen3.6-27B, Qwen3.6-35B-A3B, gemma-4-31B-it and gemma-4-26B-A4B-it models, and support per-block quantization models Qwen3.6-27B-FP8 and Qwen3.6-35B-A3B-FP8. - 2026.07 We released intel/llm-scaler-omni:0.1.0-b8 to support ComfyUI 0.27.0,more workflows and models. - 2026.07 We released intel/llm-scaler-vllm:0.21.0-b1 to support gemma-4 12B, 31B and 26B-A4B and diffusiongemma 26B-A4B models, and experimentally support XPU graph. - 2026.06 We released intel/llm-scaler-vllm:0.14.0-b8.3.2 to fix Qwen3.5/3.6-27B accuracy issues. - 2026.06 We released intel/llm-scaler-vllm:0.14.0-b8.3.1 to enable FP8 KV Cache and fix bugs for Qwen3/Qwen3.5 models. - 2026.05 We released intel/llm-scaler-vllm:0.14.0-b8.3 to improve performance for Qwen3.5/3.6 series and Qwen3-Coder-Next, and enabled model streaming load to reduce peak memory. - 2026.05 We released intel/llm-scaler-vllm:1.4 or, intel/llm-scaler-vllm:0.14.0-b8.2.1 with new platform image and support Intelยฎ Arcโ„ข Pro B70 GPU. - 2026.05 We released intel/llm-scaler-omni:0.1.0-b7 for more model workflows and performance improvments. - 2026.03 We released intel/llm-scaler-vllm:0.14.0-b8.1 to support Qwen3.5-27B, Qwen3.5-35B-A3B and Qwen3.5-122B-A10B FP8/INT4 online quantization, GPTQ - 2026.03 We released intel/llm-scaler-omni:0.1.0-b6 for ComfyUI to support CacheDiT and torch.compile , ComfyUI-GGUF, and more model workflows, and support FP8 for SGLang Diffusion. - 2026.03 We released intel/llm-scaler-vllm:0.14.0-b8 for vLLM 0.14.0 and PyTorch 2.10 support, various new models support and performance improvement. - 2026.01 We released intel/llm-scaler-vllm:1.3 or, intel/llm-scaler-vllm:0.11.1-b7 for vLLM 0.11.1 and PyTorch 2.9 support, various new models support and performance improvement. - 2026.01 We released intel/llm-scaler-omni:0.1.0-b5 for Python 3.12 and PyTorch 2.9 support, various ComfyUI workflows and more SGLang Diffusion support. - 2025.12 We released intel/llm-scaler-vllm:1.2 , same image as intel/llm-scaler-vllm:0.10.2-b6 . - 2025.12 We released intel/llm-scaler-omni:0.1.0-b4 to support ComfyUI workflows for Z-Image-Turbo, Hunyuan-Video-1.5 T2V/I2V with multi-XPU, and experimentially support SGLang Diffusion. - 2025.11 We released intel/llm-scaler-vllm:0.10.2-b6 to support Qwen3-VL Dense/MoE , Qwen3-Omni, Qwen3-30B-A3B MoE Int4 , MinerU 2.5, ERNIE-4.5-vl etc. - 2025.11 We released intel/llm-scaler-vllm:0.10.2-b5 to support gpt-oss models and released intel/llm-scaler-omni:0.1.0-b3 to support more ComfyUI workflows, and Windows installation. - 2025.10 We released intel/llm-scaler-omni:0.1.0-b2 to support more models with ComfyUI workflows and Xinference. - 2025.09 We released intel/llm-scaler-vllm:0.10.0-b3 to support more models MinerU, MiniCPM-v-4.5 etc , and released intel/llm-scaler-omni:0.1.0-b1 to enable first omni GenAI models using ComfyUI and Xinference on Arc Pro B60 GPU. - 2025.08 We released intel/llm-scaler-vllm:1.0 . llm-scaler-vllm supports running text generation models using vLLM, featuring: support P2P or USM CCL and INT4 quantized online serving, plus pre-quantized FP8 model support FP8 and Embedding model support Reranker model support Multi-Modal model support Omni , Tensor Parallel and Pipeline Parallel Data Parallel - Finding maximum Context Length - Multi-Modal WebUI - BPE-Qwen tokenizer Please follow the instructions in the Getting Started /intel/llm-scaler/blob/main/vllm/README.md/ 1-getting-started-and-usage to use llm-scaler-vllm . | Model Name | FP16 | Dynamic Online FP8 | Dynamic Online Int4 | MXFP4 | Notes | |---|---|---|---|---|---| | openai/gpt-oss-20b | โœ… | |||| | openai/gpt-oss-120b | โœ… | |||| | deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | โœ… | โœ… | โœ… | || | deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | โœ… | โœ… | โœ… | || | deepseek-ai/DeepSeek-R1-Distill-Llama-8B | โœ… | โœ… | โœ… | || | deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | โœ… | โœ… | โœ… | || | deepseek-ai/DeepSeek-R1-Distill-Qwen-32B | โœ… | โœ… | โœ… | || | deepseek-ai/DeepSeek-R1-Distill-Llama-70B | โœ… | โœ… | โœ… | || | deepseek-ai/DeepSeek-R1-0528-Qwen3-8B | โœ… | โœ… | โœ… | || | deepseek-ai/DeepSeek-V2-Lite | โœ… | โœ… | export VLLM MLA DISABLE=1 | || | deepseek-ai/deepseek-coder-33b-instruct | โœ… | โœ… | โœ… | || | Qwen/Qwen3-8B | โœ… | โœ… | โœ… | || | Qwen/Qwen3-14B | โœ… | โœ… | โœ… | || | Qwen/Qwen3-32B | โœ… | โœ… | โœ… | || | Qwen/Qwen3-30B-A3B | โœ… | โœ… | โœ… | || | Qwen/Qwen3-235B-A22B | โœ… | |||| | Qwen/Qwen3-Coder-30B-A3B-Instruct | โœ… | โœ… | โœ… | || | Qwen/Qwen3-Coder-Next | โœ… | โœ… | ||| | Qwen/Qwen3.5/3.6-27B | โœ… | โœ… | โœ… | || | Qwen/Qwen3.5/3.6-35B-A3B | โœ… | โœ… | โœ… | || | Qwen/Qwen3.6-27B-FP8 | Pre-quantized offline FP8 model | |||| | Qwen/Qwen3.6-35B-A3B-FP8 | Pre-quantized offline FP8 model | |||| | Qwen/Qwen3.5-122B-A10B | โœ… | โœ… | ||| | Qwen/QwQ-32B | โœ… | โœ… | โœ… | || | mistralai/Ministral-8B-Instruct-2410 | โœ… | โœ… | โœ… | || | mistralai/Mixtral-8x7B-Instruct-v0.1 | โœ… | โœ… | โœ… | || | meta-llama/Llama-3.1-8B | โœ… | โœ… | โœ… | || | meta-llama/Llama-3.1-70B | โœ… | โœ… | โœ… | || | baichuan-inc/Baichuan2-7B-Chat | โœ… | โœ… | โœ… | with chat template | | | baichuan-inc/Baichuan2-13B-Chat | โœ… | โœ… | โœ… | with chat template | | | THUDM/CodeGeex4-All-9B | โœ… | โœ… | โœ… | with chat template | | | zai-org/GLM-4-9B-0414 | โœ… | use bfloat16 | ||| | zai-org/GLM-4-32B-0414 | โœ… | use bfloat16 | ||| | zai-org/GLM-4.5-Air | โœ… | โœ… | ||| | zai-org/GLM-4.7-Flash | โœ… | โœ… | ||| | ByteDance-Seed/Seed-OSS-36B-Instruct | โœ… | โœ… | โœ… | || | miromind-ai/MiroThinker-v1.5-30B | โœ… | โœ… | โœ… | || | tencent/Hunyuan-0.5B-Instruct | โœ… | โœ… | โœ… | follow the guide in | here /intel/llm-scaler/blob/main/vllm/README.md 31-how-to-use-hunyuan-7b-instruct Reference Commands /intel/llm-scaler/blob/main/vllm/README.md/ 33-reference-commands-for-running-gemma-4-models-and-diffusiongemma Reference Commands /intel/llm-scaler/blob/main/vllm/README.md/ 33-reference-commands-for-running-gemma-4-models-and-diffusiongemma here /intel/llm-scaler/blob/main/vllm/README.md 32-how-to-use-paddleocr --quantization fp8 here https://github.com/vllm-project/vllm/blob/2f4226fe5280b60c47b4f6f01d9b18ac9cda2038/examples/pooling/embed/vision embedding online.py here https://github.com/vllm-project/vllm/blob/2f4226fe5280b60c47b4f6f01d9b18ac9cda2038/examples/pooling/score/vision rerank api online.py llm-scaler-omni supports running image/voice/video generation etc., featuring Omni Studio mode using ComfyUI and Omni Serving mode via SGLang Diffusion or Xinference . Please follow the instructions in the Getting Started /intel/llm-scaler/blob/main/omni/README.md/ getting-started-with-omni-docker-image to use llm-scaler-omni . | Qwen-Image | Multi B60 Wan2.2-T2V-14B | |---|---| Omni Stuido supports Image Generation/Edit, Video Generation, Audio Generation, 3D Generation etc. | Model Category | Model | Type | |---|---|---| Image Generation | Qwen-Image, Qwen-Image-Edit | Text-to-Image, Image Editing | Image Generation | Stable Diffusion 3.5 | Text-to-Image, ControlNet | Image Generation | Z-Image-Turbo | Text-to-Image | Image Generation | Flux.1, Flux.1 Kontext dev | Text-to-Image, Multi-Image Reference, ControlNet | Image Generation | FireRed-Image-Edit-1.1 | Image Editing | Video Generation | Wan2.2 TI2V 5B, Wan2.2 T2V 14B, Wan2.2 I2V 14B | Text-to-Video, Image-to-Video | Video Generation | Wan2.2 Animate 14B | Video Animation | Video Generation | HunyuanVideo 1.5 8.3B | Text-to-Video, Image-to-Video | Video Generation | LTX-2 | Text-to-Video, Image-to-Video | 3D Generation | Hunyuan3D 2.1 | Text/Image-to-3D | Audio Generation | VoxCPM1.5, IndexTTS 2 | Text-to-Speech, Voice Cloning | Video Upscaling | SeedVR2 | Video Restoration and Upscaling | Please check ComfyUI Support /intel/llm-scaler/blob/main/omni/README.md/ comfyui for more details. Omni Serving supports Image Generation, Audio Generation etc. - Image Generation /v1/images/generations : Stable Diffusion 3.5, Flux.1-dev - Text to Speech /v1/audio/speech : Kokoro 82M - Speech to Text /v1/audio/transcriptions : whisper-large-v3 Please check Xinference Support /intel/llm-scaler/blob/main/omni/README.md/ xinference for more details. - Please check out the Docker image releases for llm-scaler-vllm /intel/llm-scaler/blob/main/Releases.md/ llm-scaler-vllm and llm-scaler-omni /intel/llm-scaler/blob/main/Releases.md/ llm-scaler-omni - Please report a bug or raise a feature request by opening a Github Issue https://github.com/intel/llm-scaler/issues