{"slug": "low-rank-prompt-learning-for-vision-language-models-with-fixed-token-bases", "title": "Low-Rank Prompt Learning for Vision-Language Models with Fixed-Token Bases", "summary": "A new arXiv paper (2609.09462v1) reports that factorizing the dense CoOp prompt matrix P ∈ R^(m×d) as P = BA cuts trainable prompt parameters from md to r(m+d), and to rd once the token-side factor B is fixed, matching or improving dense CoOp across seven few-shot benchmarks and two CLIP backbones. The authors find the token-side factor B need not be learned at all: fixing B to a Gaussian, orthogonal, SVD-derived, or even random basis and training only the embedding-side factor A stays on par with the fully trainable factorization, and a source-trained B offers no advantage over a random one. A prompt-factor asymmetry and a local update-space dimension gap explain why fixing B is far less restrictive than fixing A, with a smoothness-only guarantee certifying convergence when optimizing A over a fixed B.", "body_md": "arXiv:2609.09462v1 Announce Type: new \nAbstract: Prompt learning adapts CLIP to downstream recognition by replacing hand-written templates with learned continuous context vectors, which in Context Optimization (CoOp) form a dense prompt matrix $\\mathbf{P}\\in\\mathbb{R}^{m\\times d}$ trained from only a few examples per class. We study whether this matrix is over-parameterized by factorizing it as $\\mathbf{P}=\\mathbf{B}\\mathbf{A}$, which cuts the trainable prompt parameters from $md$ to $r(m+d)$, and to $rd$ once the token-side factor $\\mathbf{B}$ is fixed. Across seven few-shot benchmarks and two CLIP backbones, low-rank prompts match or improve dense CoOp at far fewer parameters, with the clearest gains on low-shot base-to-new generalization. We then find that the token-side factor need not be learned at all: fixing $\\mathbf{B}$ to a Gaussian, orthogonal, SVD-derived, or even random basis and training only the embedding-side factor $\\mathbf{A}$ stays on par with the fully trainable factorization, and a source-trained $\\mathbf{B}$ offers no advantage over a random one. A prompt-factor asymmetry and a local update-space dimension gap show why fixing $\\mathbf{B}$ is far less restrictive than fixing $\\mathbf{A}$, and a smoothness-only guarantee certifies that optimizing $\\mathbf{A}$ over a fixed $\\mathbf{B}$ converges. In the CLIP prompt setting, the embedding-side coefficients carry the adaptation while the token basis can simply be fixed.", "url": "https://wpnews.pro/news/low-rank-prompt-learning-for-vision-language-models-with-fixed-token-bases", "canonical_source": "https://arxiv.org/abs/2609.09462", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 04:24:37.587351+00:00", "lang": "en", "topics": ["computer-vision", "machine-learning", "ai-research", "natural-language-processing"], "entities": ["CLIP", "Context Optimization (CoOp)", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/low-rank-prompt-learning-for-vision-language-models-with-fixed-token-bases", "markdown": "https://wpnews.pro/news/low-rank-prompt-learning-for-vision-language-models-with-fixed-token-bases.md", "text": "https://wpnews.pro/news/low-rank-prompt-learning-for-vision-language-models-with-fixed-token-bases.txt", "jsonld": "https://wpnews.pro/news/low-rank-prompt-learning-for-vision-language-models-with-fixed-token-bases.jsonld"}}