Is there any full Vision transformer model for object detection? Multiple full Vision Transformer (ViT) object detection models exist, including LW-DETR, RF-DETR, YOLOS, ViTDet, and WB-DETR, each differing in architectural purity and usability. The choice depends on whether the user prioritizes a plain ViT backbone with a DETR-style decoder, practical performance, or minimal modifications to vanilla ViT. Probably, yes: If by “full Vision Transformer” you mean an object detector whose main visual backbone is a ViT rather than a CNN , there are several examples now. For the specific shape you mentioned — something like a plain ViT backbone + a DETR-style head/decoder — I would start with LW-DETR . Hugging Face describes its architecture as a If you want a newer, practical detector built around a pretrained ViT backbone, RF-DETR is also worth looking at. It uses a And if what you mean is closer to “vanilla ViT itself turned into an object detector” , then YOLOS is probably the cleanest reference point. It modifies ViT only minimally and adds detection tokens/heads rather than attaching a conventional DETR encoder-decoder stack. So, roughly: | What you are looking for | Candidate | Transformers usability | |---|---|---| | plain ViT + DETR-style decoder | | pipeline / AutoModelForObjectDetection pipeline / AutoModelForObjectDetection The only thing I would be careful about is the word “entirely” , because there are a few different meanings hiding inside it. If I were choosing where to start: Want the closest thing to "plain ViT backbone + DETR head"? → LW-DETR Want a newer practical ViT-based detector with good current HF support? → RF-DETR Want the cleanest vanilla-ViT-as-detector reference? → YOLOS Want to build a ViT backbone + strong DETR-family detector yourself? → ViTDet + DINO/detrex Want specifically "no CNN backbone" as a research constraint? → WB-DETR Want literally zero Conv2d / convolution operations? → treat that as a separate requirement and audit the exact model implementation So the answer is definitely not “there is no full-ViT object detector”; rather, there are several families already. The main choice is what you mean by “full” and whether your priority is architectural purity, a DETR-like component boundary, current Transformers usability, or practical detection performance.