ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder, achieved 75.41% average task success and an 8.20% average obstacle collision rate across 61 real-robot greenhouse trials per method, totaling 366 executions, according to the arXiv paper 2609.10918v1. The framework extracts a structured target-obstacle-background representation so a downstream alignment policy can generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. ObstaDiff outperformed representative imitation-learning baselines and improved generalization in cluttered agricultural scenes. arXiv:2609.10918v1 Announce Type: cross Abstract: Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-world scenes with unstructured obstacles remains a key generalization challenge. We present ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder. ObstaDiff extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. We evaluate ObstaDiff on 61 real-robot greenhouse trials per method 366 executions in total . ObstaDiff achieves 75.41% average task success and 8.20% average obstacle collision rate, outperforming representative imitation-learning baselines and improving generalization in cluttered agricultural scenes.