{"slug": "tape-ml-a-compact-structured-representation-for-multi-task-computer-vision", "title": "TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision", "summary": "Researchers introduced TAPe+ML v3, a compact multi-task computer vision system built on the TAPe (Theory of Active Perception) structured representation, using fewer than 100,000 parameters. The system reports 84.7 mAP50 and 65.3 mAP50-95 on COCO object detection, 80.7 mask mAP50 and 58.4 mask mAP50-95 on COCO instance segmentation, 92 percent validation accuracy on Imagenette versus a raw-pixel baseline under identical training, and 89.9 percent Top-1 accuracy on ImageNet-Real. The authors state the results suggest shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.", "body_md": "arXiv:2609.20869v1 Announce Type: new \nAbstract: We present TAPe+ML v3, a compact computer vision system based on TAPe (Theory of Active Perception), a structured representation that encodes relations among perceptual elements before recognition. Instead of operating directly on pixel tensors, the system uses a shared TAPe representation and a modular recognition architecture for image classification, object detection, and instance segmentation.\n  TAPe+ML v3 combines background and contour processing, local object localization, prototype-based classification, and a coordinator for specialized submodels. Across the reported experiments, it uses fewer than 100,000 parameters. On COCO object detection, it obtains 84.7 mAP50 and 65.3 mAP50-95. On COCO instance segmentation, it obtains 80.7 mask mAP50 and 58.4 mask mAP50-95. In classification experiments, it reaches 92 percent validation accuracy on Imagenette under an identical-training comparison with a raw-pixel baseline, and 89.9 percent Top-1 accuracy on ImageNet-Real. We also evaluate compactness in video scene detection and adaptation under distribution shift in an industrial pilot. The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.", "url": "https://wpnews.pro/news/tape-ml-a-compact-structured-representation-for-multi-task-computer-vision", "canonical_source": "https://arxiv.org/abs/2609.20869", "published_at": "2026-09-21 04:00:00+00:00", "updated_at": "2026-09-21 04:26:44.697269+00:00", "lang": "en", "topics": ["computer-vision", "machine-learning", "ai-research"], "entities": ["TAPe+ML v3", "TAPe (Theory of Active Perception)", "COCO", "Imagenette", "ImageNet-Real", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/tape-ml-a-compact-structured-representation-for-multi-task-computer-vision", "markdown": "https://wpnews.pro/news/tape-ml-a-compact-structured-representation-for-multi-task-computer-vision.md", "text": "https://wpnews.pro/news/tape-ml-a-compact-structured-representation-for-multi-task-computer-vision.txt", "jsonld": "https://wpnews.pro/news/tape-ml-a-compact-structured-representation-for-multi-task-computer-vision.jsonld"}}