04:00
2026-09-16
arxiv.org
artificial-intelligence
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Researchers introduced FLAT (Flexible-Length Aligned Transmodal representations), a pre-training framework that jointly optimizes a shared multimodal encoder with text-to-image and image-to-text decodβ¦