{"slug": "llm-scaler-llm-support-for-intel-s-arc-pro-b60-and-b70-gpus", "title": "LLM Scaler – LLM Support for Intel's Arc Pro B60 and B70 GPUs", "summary": "Intel has released LLM Scaler, a GenAI solution for text, image, and video generation optimized for Intel Arc Pro B60 and B70 GPUs, with the latest version intel/llm-scaler-vllm:0.21.0-b2 adding Multi-token Prediction and Lora Serving for Qwen3.6-27B, Qwen3.6-35B-A3B, gemma-4-31B-it, and gemma-4-26B-A4B-it models, plus FP8 per-block quantization support. The solution leverages frameworks such as vLLM, ComfyUI, SGLang Diffusion, and Xinference to deliver performance for state-of-the-art GenAI models.", "body_md": "LLM Scaler is an GenAI solution for text generation, image generation, video generation etc. running on Intel® Arc™ Pro B60 and B70 GPUs. LLM Scalar leverages standard frameworks such as vLLM, ComfyUI, SGLang Diffusion, Xinference etc and ensures the best performance for State-of-Art GenAI models running on Arc Pro B60/B70 GPUs.\n\n- 🔥[2026.08] We released\n`intel/llm-scaler-vllm:0.21.0-b2`\n\nto support Multi-token Prediction (MTP) and Lora Serving for Qwen3.6-27B, Qwen3.6-35B-A3B, gemma-4-31B-it and gemma-4-26B-A4B-it models, and support per-block quantization models Qwen3.6-27B-FP8 and Qwen3.6-35B-A3B-FP8. - [2026.07] We released\n`intel/llm-scaler-omni:0.1.0-b8`\n\nto support ComfyUI 0.27.0,more workflows and models. - [2026.07] We released\n`intel/llm-scaler-vllm:0.21.0-b1`\n\nto support gemma-4 (12B, 31B and 26B-A4B) and diffusiongemma (26B-A4B) models, and experimentally support XPU graph. - [2026.06] We released\n`intel/llm-scaler-vllm:0.14.0-b8.3.2`\n\nto fix Qwen3.5/3.6-27B accuracy issues. - [2026.06] We released\n`intel/llm-scaler-vllm:0.14.0-b8.3.1`\n\nto enable FP8 KV Cache and fix bugs for Qwen3/Qwen3.5 models. - [2026.05] We released\n`intel/llm-scaler-vllm:0.14.0-b8.3`\n\nto improve performance for Qwen3.5/3.6 series and Qwen3-Coder-Next, and enabled model streaming load to reduce peak memory. - [2026.05] We released\n`intel/llm-scaler-vllm:1.4`\n\n(or,`intel/llm-scaler-vllm:0.14.0-b8.2.1`\n\n) with new platform image and support Intel® Arc™ Pro B70 GPU. - [2026.05] We released\n`intel/llm-scaler-omni:0.1.0-b7`\n\nfor more model workflows and performance improvments. - [2026.03] We released\n`intel/llm-scaler-vllm:0.14.0-b8.1`\n\nto support Qwen3.5-27B, Qwen3.5-35B-A3B and Qwen3.5-122B-A10B (FP8/INT4 online quantization, GPTQ) - [2026.03] We released\n`intel/llm-scaler-omni:0.1.0-b6`\n\nfor ComfyUI to support CacheDiT and torch.compile(), ComfyUI-GGUF, and more model workflows, and support FP8 for SGLang Diffusion. - [2026.03] We released\n`intel/llm-scaler-vllm:0.14.0-b8`\n\nfor vLLM 0.14.0 and PyTorch 2.10 support, various new models support and performance improvement. - [2026.01] We released\n`intel/llm-scaler-vllm:1.3`\n\n(or,`intel/llm-scaler-vllm:0.11.1-b7`\n\n) for vLLM 0.11.1 and PyTorch 2.9 support, various new models support and performance improvement. - [2026.01] We released\n`intel/llm-scaler-omni:0.1.0-b5`\n\nfor Python 3.12 and PyTorch 2.9 support, various ComfyUI workflows and more SGLang Diffusion support. - [2025.12] We released\n`intel/llm-scaler-vllm:1.2`\n\n, same image as`intel/llm-scaler-vllm:0.10.2-b6`\n\n. - [2025.12] We released\n`intel/llm-scaler-omni:0.1.0-b4`\n\nto support ComfyUI workflows for Z-Image-Turbo, Hunyuan-Video-1.5 T2V/I2V with multi-XPU, and experimentially support SGLang Diffusion. - [2025.11] We released\n`intel/llm-scaler-vllm:0.10.2-b6`\n\nto support Qwen3-VL (Dense/MoE), Qwen3-Omni, Qwen3-30B-A3B (MoE Int4), MinerU 2.5, ERNIE-4.5-vl etc. - [2025.11] We released\n`intel/llm-scaler-vllm:0.10.2-b5`\n\nto support gpt-oss models and released`intel/llm-scaler-omni:0.1.0-b3`\n\nto support more ComfyUI workflows, and Windows installation. - [2025.10] We released\n`intel/llm-scaler-omni:0.1.0-b2`\n\nto support more models with ComfyUI workflows and Xinference. - [2025.09] We released\n`intel/llm-scaler-vllm:0.10.0-b3`\n\nto support more models (MinerU, MiniCPM-v-4.5 etc), and released`intel/llm-scaler-omni:0.1.0-b1`\n\nto enable first omni GenAI models using ComfyUI and Xinference on Arc Pro B60 GPU. - [2025.08] We released\n`intel/llm-scaler-vllm:1.0`\n\n.\n\n`llm-scaler-vllm`\n\nsupports running text generation models using vLLM, featuring:\n\nsupport (P2P or USM)**CCL** and**INT4** quantized online serving, plus pre-quantized FP8 model support**FP8** and**Embedding** model support**Reranker** model support**Multi-Modal** model support**Omni**,** Tensor Parallel**and** Pipeline Parallel****Data Parallel**- Finding maximum Context Length\n- Multi-Modal WebUI\n- BPE-Qwen tokenizer\n\nPlease follow the instructions in the [Getting Started](/intel/llm-scaler/blob/main/vllm/README.md/#1-getting-started-and-usage) to use `llm-scaler-vllm`\n\n.\n\n| Model Name | FP16 | Dynamic Online FP8 | Dynamic Online Int4 | MXFP4 | Notes |\n|---|---|---|---|---|---|\n| openai/gpt-oss-20b | ✅ | ||||\n| openai/gpt-oss-120b | ✅ | ||||\n| deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | ✅ | ✅ | ✅ | ||\n| deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | ✅ | ✅ | ✅ | ||\n| deepseek-ai/DeepSeek-R1-Distill-Llama-8B | ✅ | ✅ | ✅ | ||\n| deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | ✅ | ✅ | ✅ | ||\n| deepseek-ai/DeepSeek-R1-Distill-Qwen-32B | ✅ | ✅ | ✅ | ||\n| deepseek-ai/DeepSeek-R1-Distill-Llama-70B | ✅ | ✅ | ✅ | ||\n| deepseek-ai/DeepSeek-R1-0528-Qwen3-8B | ✅ | ✅ | ✅ | ||\n| deepseek-ai/DeepSeek-V2-Lite | ✅ | ✅ | export VLLM_MLA_DISABLE=1 | ||\n| deepseek-ai/deepseek-coder-33b-instruct | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3-8B | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3-14B | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3-32B | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3-30B-A3B | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3-235B-A22B | ✅ | ||||\n| Qwen/Qwen3-Coder-30B-A3B-Instruct | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3-Coder-Next | ✅ | ✅ | |||\n| Qwen/Qwen3.5/3.6-27B | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3.5/3.6-35B-A3B | ✅ | ✅ | ✅ | ||\n| Qwen/Qwen3.6-27B-FP8 | Pre-quantized offline FP8 model | ||||\n| Qwen/Qwen3.6-35B-A3B-FP8 | Pre-quantized offline FP8 model | ||||\n| Qwen/Qwen3.5-122B-A10B | ✅ | ✅ | |||\n| Qwen/QwQ-32B | ✅ | ✅ | ✅ | ||\n| mistralai/Ministral-8B-Instruct-2410 | ✅ | ✅ | ✅ | ||\n| mistralai/Mixtral-8x7B-Instruct-v0.1 | ✅ | ✅ | ✅ | ||\n| meta-llama/Llama-3.1-8B | ✅ | ✅ | ✅ | ||\n| meta-llama/Llama-3.1-70B | ✅ | ✅ | ✅ | ||\n| baichuan-inc/Baichuan2-7B-Chat | ✅ | ✅ | ✅ | with chat_template | |\n| baichuan-inc/Baichuan2-13B-Chat | ✅ | ✅ | ✅ | with chat_template | |\n| THUDM/CodeGeex4-All-9B | ✅ | ✅ | ✅ | with chat_template | |\n| zai-org/GLM-4-9B-0414 | ✅ | use bfloat16 | |||\n| zai-org/GLM-4-32B-0414 | ✅ | use bfloat16 | |||\n| zai-org/GLM-4.5-Air | ✅ | ✅ | |||\n| zai-org/GLM-4.7-Flash | ✅ | ✅ | |||\n| ByteDance-Seed/Seed-OSS-36B-Instruct | ✅ | ✅ | ✅ | ||\n| miromind-ai/MiroThinker-v1.5-30B | ✅ | ✅ | ✅ | ||\n| tencent/Hunyuan-0.5B-Instruct | ✅ | ✅ | ✅ | follow the guide in\n|\n\n[here](/intel/llm-scaler/blob/main/vllm/README.md#31-how-to-use-hunyuan-7b-instruct)[Reference Commands](/intel/llm-scaler/blob/main/vllm/README.md/#33-reference-commands-for-running-gemma-4-models-and-diffusiongemma)[Reference Commands](/intel/llm-scaler/blob/main/vllm/README.md/#33-reference-commands-for-running-gemma-4-models-and-diffusiongemma)[here](/intel/llm-scaler/blob/main/vllm/README.md#32-how-to-use-paddleocr)`--quantization fp8`\n\n[here](https://github.com/vllm-project/vllm/blob/2f4226fe5280b60c47b4f6f01d9b18ac9cda2038/examples/pooling/embed/vision_embedding_online.py)[here](https://github.com/vllm-project/vllm/blob/2f4226fe5280b60c47b4f6f01d9b18ac9cda2038/examples/pooling/score/vision_rerank_api_online.py)`llm-scaler-omni`\n\nsupports running image/voice/video generation etc., featuring `Omni Studio`\n\nmode (using ComfyUI) and `Omni Serving`\n\nmode (via SGLang Diffusion or Xinference).\n\nPlease follow the instructions in the [Getting Started](/intel/llm-scaler/blob/main/omni/README.md/#getting-started-with-omni-docker-image) to use `llm-scaler-omni`\n\n.\n\n| Qwen-Image | Multi B60 Wan2.2-T2V-14B |\n|---|---|\n\n`Omni Stuido`\n\nsupports Image Generation/Edit, Video Generation, Audio Generation, 3D Generation etc.\n\n| Model Category | Model | Type |\n|---|---|---|\nImage Generation |\nQwen-Image, Qwen-Image-Edit | Text-to-Image, Image Editing |\nImage Generation |\nStable Diffusion 3.5 | Text-to-Image, ControlNet |\nImage Generation |\nZ-Image-Turbo | Text-to-Image |\nImage Generation |\nFlux.1, Flux.1 Kontext dev | Text-to-Image, Multi-Image Reference, ControlNet |\nImage Generation |\nFireRed-Image-Edit-1.1 | Image Editing |\nVideo Generation |\nWan2.2 TI2V 5B, Wan2.2 T2V 14B, Wan2.2 I2V 14B | Text-to-Video, Image-to-Video |\nVideo Generation |\nWan2.2 Animate 14B | Video Animation |\nVideo Generation |\nHunyuanVideo 1.5 8.3B | Text-to-Video, Image-to-Video |\nVideo Generation |\nLTX-2 | Text-to-Video, Image-to-Video |\n3D Generation |\nHunyuan3D 2.1 | Text/Image-to-3D |\nAudio Generation |\nVoxCPM1.5, IndexTTS 2 | Text-to-Speech, Voice Cloning |\nVideo Upscaling |\nSeedVR2 | Video Restoration and Upscaling |\n\nPlease check [ComfyUI Support](/intel/llm-scaler/blob/main/omni/README.md/#comfyui) for more details.\n\n`Omni Serving`\n\nsupports Image Generation, Audio Generation etc.\n\n- Image Generation (\n`/v1/images/generations`\n\n): Stable Diffusion 3.5, Flux.1-dev - Text to Speech (\n`/v1/audio/speech`\n\n): Kokoro 82M - Speech to Text (\n`/v1/audio/transcriptions`\n\n): whisper-large-v3\n\nPlease check [Xinference Support](/intel/llm-scaler/blob/main/omni/README.md/#xinference) for more details.\n\n- Please check out the Docker image releases for\n[llm-scaler-vllm](/intel/llm-scaler/blob/main/Releases.md/#llm-scaler-vllm)and[llm-scaler-omni](/intel/llm-scaler/blob/main/Releases.md/#llm-scaler-omni)\n\n- Please report a bug or raise a feature request by opening a\n[Github Issue](https://github.com/intel/llm-scaler/issues)", "url": "https://wpnews.pro/news/llm-scaler-llm-support-for-intel-s-arc-pro-b60-and-b70-gpus", "canonical_source": "https://github.com/intel/llm-scaler", "published_at": "2026-08-09 19:17:00+00:00", "updated_at": "2026-08-09 19:35:05.952787+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-infrastructure", "ai-products"], "entities": ["Intel", "Intel Arc Pro B60", "Intel Arc Pro B70", "vLLM", "ComfyUI", "SGLang Diffusion", "Xinference", "Qwen3.6-27B"], "alternates": {"html": "https://wpnews.pro/news/llm-scaler-llm-support-for-intel-s-arc-pro-b60-and-b70-gpus", "markdown": "https://wpnews.pro/news/llm-scaler-llm-support-for-intel-s-arc-pro-b60-and-b70-gpus.md", "text": "https://wpnews.pro/news/llm-scaler-llm-support-for-intel-s-arc-pro-b60-and-b70-gpus.txt", "jsonld": "https://wpnews.pro/news/llm-scaler-llm-support-for-intel-s-arc-pro-b60-and-b70-gpus.jsonld"}}