cd /news/computer-vision/tape-ml-a-compact-structured-represe… · home topics computer-vision article
[ARTICLE · art-135565] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

Researchers introduced TAPe+ML v3, a compact multi-task computer vision system built on the TAPe (Theory of Active Perception) structured representation, using fewer than 100,000 parameters. The system reports 84.7 mAP50 and 65.3 mAP50-95 on COCO object detection, 80.7 mask mAP50 and 58.4 mask mAP50-95 on COCO instance segmentation, 92 percent validation accuracy on Imagenette versus a raw-pixel baseline under identical training, and 89.9 percent Top-1 accuracy on ImageNet-Real. The authors state the results suggest shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.

read1 min views1 publishedSep 21, 2026

arXiv:2609.20869v1 Announce Type: new Abstract: We present TAPe+ML v3, a compact computer vision system based on TAPe (Theory of Active Perception), a structured representation that encodes relations among perceptual elements before recognition. Instead of operating directly on pixel tensors, the system uses a shared TAPe representation and a modular recognition architecture for image classification, object detection, and instance segmentation. TAPe+ML v3 combines background and contour processing, local object localization, prototype-based classification, and a coordinator for specialized submodels. Across the reported experiments, it uses fewer than 100,000 parameters. On COCO object detection, it obtains 84.7 mAP50 and 65.3 mAP50-95. On COCO instance segmentation, it obtains 80.7 mask mAP50 and 58.4 mask mAP50-95. In classification experiments, it reaches 92 percent validation accuracy on Imagenette under an identical-training comparison with a raw-pixel baseline, and 89.9 percent Top-1 accuracy on ImageNet-Real. We also evaluate compactness in video scene detection and adaptation under distribution shift in an industrial pilot. The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.

── more in #computer-vision 4 stories · sorted by recency
── more on @tape+ml v3 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tape-ml-a-compact-st…] indexed:0 read:1min 2026-09-21 ·