cd /news/artificial-intelligence/pistis-technical-report · home › topics › artificial-intelligence › article
[ARTICLE · art-139441] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Pistis Technical Report

The Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, was introduced in arXiv paper 2609.28554v1 using a post-training framework that combines large-scale multimodal supervised fine-tuning with Interleaved Distillation and Reinforcement Learning (IDRL). IDRL alternates on-policy distillation and reinforcement learning in a single training loop, yielding Pistis-Thinking for deep multimodal reasoning and Pistis-Agentic for long-horizon planning, iterative reasoning and tool use, with both scales outperforming their base models. The authors also present Pistis-Auto-Harnessing (PAH), which improves agent inference-harness performance without updating model parameters or increasing the interaction budget.

by read1 min views1 publishedSep 25, 2026

arXiv:2609.28554v1 Announce Type: new Abstract: We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade-offs. At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search. Both scales outperform their corresponding base models. Beyond model-parameter optimization, we further introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @pistis 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/pistis-technical-rep…] indexed:0 read:1min 2026-09-25 · —