cd /news/artificial-intelligence/flow3d-opd-multi-teacher-on-policy-d… · home topics artificial-intelligence article
[ARTICLE · art-125489] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer

Researchers posted arXiv:2609.07137v1, a paper introducing Flow3D-OPD, a two-stage post-training framework that applies multi-teacher on-policy distillation to image-to-3D generation built on flow-matching diffusion Transformers. The framework first uses a semi-policy to strengthen the pretrained model and an agentic verifier for 3D geometric quality, then cultivates domain-specialized teacher models via direct preference optimization before consolidating them into a unified student model through on-policy distillation with hard task-routing sampling and gradient accumulation. The authors report consistent improvements across all geometric quality dimensions, with the student surpassing all teacher models on the average metric.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.07137v1 Announce Type: cross Abstract: Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training strategy remains largely unexplored. There exist several critical bottlenecks in reinforcement learning: the inherent difficulty of defining comprehensive rewards for 3D geometric quality, and the gradient interference that arises when jointly optimizing heterogeneous objectives. Inspired by the practicability of on-policy distillation (OPD) in large language models and image generation, we propose \textbf{Flow3D-OPD}, a two-stage post-training framework that introduces multi-teacher distillation into 3D geometry generation. In the first stage, we utilize the semi-policy to enhance the foundational capability of the pretrained model and then design an agentic verifier for 3D geometric quality evaluation. Based on the verifier, we could cultivate domain-specialized teacher models via direct preference optimization (DPO). In the second stage, we consolidate heterogeneous expertise into a unified student model through on-policy distillation with hard task-routing sampling and gradient accumulation, which could mitigate the gradient interference in joint optimization. Without relying on elaborate modifications, our straightforward yet effective design achieves consistent improvements across all geometric quality dimensions and surpasses all teacher models in the average metric. Extensive experiments demonstrate that our approach provides an effective paradigm for reinforcement learning in 3D generation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @flow3d-opd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/flow3d-opd-multi-tea…] indexed:0 read:1min 2026-09-10 ·