{"slug": "pistis-technical-report", "title": "Pistis Technical Report", "summary": "The Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, was introduced in arXiv paper 2609.28554v1 using a post-training framework that combines large-scale multimodal supervised fine-tuning with Interleaved Distillation and Reinforcement Learning (IDRL). IDRL alternates on-policy distillation and reinforcement learning in a single training loop, yielding Pistis-Thinking for deep multimodal reasoning and Pistis-Agentic for long-horizon planning, iterative reasoning and tool use, with both scales outperforming their base models. The authors also present Pistis-Auto-Harnessing (PAH), which improves agent inference-harness performance without updating model parameters or increasing the interaction budget.", "body_md": "arXiv:2609.28554v1 Announce Type: new \nAbstract: We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.", "url": "https://wpnews.pro/news/pistis-technical-report", "canonical_source": "https://arxiv.org/abs/2609.28554", "published_at": "2026-09-25 04:00:00+00:00", "updated_at": "2026-09-25 04:29:14.628193+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-agents", "ai-research"], "entities": ["Pistis", "Qwen3.6", "Qwen3.5", "Pistis-Thinking", "Pistis-Agentic", "Pistis-Auto-Harnessing", "IDRL", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/pistis-technical-report", "markdown": "https://wpnews.pro/news/pistis-technical-report.md", "text": "https://wpnews.pro/news/pistis-technical-report.txt", "jsonld": "https://wpnews.pro/news/pistis-technical-report.jsonld"}}