cd /news/computer-vision/dart-depth-as-target-pretraining-for… · home topics computer-vision article
[ARTICLE · art-121885] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

Researchers introduced DART, an RGB-D pretraining method that builds on DINOv2 by adding a pixel-space depth reconstruction objective supervised by pseudo-labeled depth, improving vision foundation models for surgery. Across eight surgical benchmarks, DART outperformed natural-image and in-domain baselines, including vanilla DINOv2, in dense prediction and image-level understanding, with depth proving more effective than Canny edges as a target. The method uses depth only during pretraining, keeping fine-tuning and inference RGB-only, and requires no extra labels or added inference cost.

read1 min views1 publishedSep 7, 2026

arXiv:2609.04555v1 Announce Type: new Abstract: Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained backbone can provide rich representations for many downstream tasks. Yet the dominant self-supervised pretraining paradigm uses only RGB images, leaving readily available complementary signals, such as depth maps, unused. This is a particular missed opportunity in surgery, where natural-image VFMs transfer poorly while the scene geometry is rich and informative. With strong off-the-shelf models now able to produce pseudo-labeled dense depth for any image corpus, we hypothesize that such signals can be folded into pretraining to learn better representations. We present DART, an RGB-D pretraining recipe that builds on DINOv2 with a simple modification: a pixel-space depth reconstruction objective applied to masked iBOT patches, supervised by pseudo-labeled depth. Depth is used only during pretraining, so fine-tuning and inference remain RGB-only. We find that this pixel-level reconstruction head improves representation quality rather than disrupting it. We further show that depth, which encodes scene geometry, is more effective as a target than alternative dense signals such as Canny edges, confirming that the gains stem from depth rather than added supervision alone. Across eight surgical benchmarks spanning segmentation, depth estimation, and image-level recognition, DART outperforms both natural-image and in-domain baselines, including a vanilla DINOv2 trained on identical data, improving dense prediction while also strengthening image-level understanding. More broadly, DART shows that freely available geometric pseudo-labels can strengthen foundation model pretraining without extra labels or added inference cost, pointing toward stronger backbones for surgery.

── more in #computer-vision 4 stories · sorted by recency
── more on @dart 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dart-depth-as-target…] indexed:0 read:1min 2026-09-07 ·