cd /news/artificial-intelligence/foldingagent-inferring-parametric-or… · home topics artificial-intelligence article
[ARTICLE · art-118517] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

Researchers introduced FoldingAgent, an agentic framework that infers parametric folding programs from origami demonstration videos using a vision-language model and specialized tools, achieving executable and physically plausible procedures on the new PurelandFold benchmark. The framework mitigates compounding errors by sequentially re-planning actions, bridging unstructured visual demonstrations and structured computational representations.

read1 min views3 publishedSep 2, 2026

arXiv:2609.00377v1 Announce Type: new Abstract: We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @foldingagent 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/foldingagent-inferri…] indexed:0 read:1min 2026-09-02 ·