{"slug": "efficient-tracking-and-understanding-object-transformations", "title": "Efficient Tracking and Understanding Object Transformations", "summary": "Researchers propose FluxGraph, a reactive variant of TubeletGraph, that uses SAM2's internal multi-mask disagreement as a lightweight trigger for transformation detection, achieving ~3.3× speedup on VOST while improving tracking performance and preserving state graph quality. FluxGraph also shows consistent speedups of 3.7–10.7× across VSCOS, M³-VOS, and DAVIS17, addressing the high inference cost of TubeletGraph (~$4.4 seconds per object-frame on VOST) that precludes real-time deployment.", "body_md": "arXiv:2607.19743v1 Announce Type: new\nAbstract: Tracking objects through state transformations is essential for understanding real-world dynamics. However, existing methods are computationally expensive. TubeletGraph recently showed impressive capabilities, but its inference cost (~$4.4$ seconds per object-frame on VOST) precludes any real-time deployment possibilities. We observe that TubeletGraph's overhead arises from building a spatiotemporal partition of the input video: (1) entity segmentation is computed densely for every frame regardless of whether a transformation occurs, and (2) every entity in the scene is tracked, scaling cost with scene complexity rather than the number of transformations of interest. To address both, we propose FluxGraph, a reactive variant that uses SAM2's internal multi-mask disagreement as a lightweight trigger for transformation detection, and removes the need for tracking all entities in the given video. FluxGraph is ~$3.3\\times$ faster than TubeletGraph on VOST while improving tracking performance and preserving state graph quality. Furthermore, we also observe consistent speedups of $3.7-10.7\\times$ across VSCOS, M$^3$-VOS, and DAVIS17 while maintaining performance. Code is publicly available at https://github.com/YihongSun/FluxGraph.", "url": "https://wpnews.pro/news/efficient-tracking-and-understanding-object-transformations", "canonical_source": "https://arxiv.org/abs/2607.19743", "published_at": "2026-07-23 04:00:00+00:00", "updated_at": "2026-07-23 04:26:17.766009+00:00", "lang": "en", "topics": ["computer-vision", "machine-learning"], "entities": ["FluxGraph", "TubeletGraph", "SAM2", "VOST", "VSCOS", "M³-VOS", "DAVIS17", "YihongSun"], "alternates": {"html": "https://wpnews.pro/news/efficient-tracking-and-understanding-object-transformations", "markdown": "https://wpnews.pro/news/efficient-tracking-and-understanding-object-transformations.md", "text": "https://wpnews.pro/news/efficient-tracking-and-understanding-object-transformations.txt", "jsonld": "https://wpnews.pro/news/efficient-tracking-and-understanding-object-transformations.jsonld"}}