{"slug": "vector-symbolic-policy-gradient", "title": "Vector Symbolic Policy Gradient", "summary": "Researchers introduced Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to the encoded state. They proved that under the standard softmax policy-gradient surrogate, its update is exactly advantage-weighted hypervector bundling followed by normalization, and that each trained action hypervector is a fixed-size compressed kernel memory. For bipolar action memories, greedy action selection is stable under random bit flips, with failure probability decaying exponentially in the hypervector dimension.", "body_md": "arXiv:2608.18404v1 Announce Type: new\nAbstract: We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to the encoded state. Under the standard softmax policy-gradient surrogate, we prove that its update is exactly advantage-weighted hypervector bundling followed by normalization, and therefore supports standard advantage estimators. We further show that each trained action hypervector is a fixed-size compressed kernel memory, storing an advantage-weighted kernel expansion over visited states and transferring evidence according to the encoder-induced similarity. This provides a concrete mechanism that can support sample-efficient learning without increasing inference-time memory. Finally, for bipolar action memories, we prove that greedy action selection is stable under random bit flips, with failure probability decaying exponentially in the hypervector dimension. VSPG thus connects VSA action memories, log-linear policy gradients, and kernel policy search while providing a quantitative robustness guarantee.", "url": "https://wpnews.pro/news/vector-symbolic-policy-gradient", "canonical_source": "https://arxiv.org/abs/2608.18404", "published_at": "2026-08-20 04:00:00+00:00", "updated_at": "2026-08-20 04:14:40.228076+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["Vector-Symbolic Policy Gradient", "VSPG"], "alternates": {"html": "https://wpnews.pro/news/vector-symbolic-policy-gradient", "markdown": "https://wpnews.pro/news/vector-symbolic-policy-gradient.md", "text": "https://wpnews.pro/news/vector-symbolic-policy-gradient.txt", "jsonld": "https://wpnews.pro/news/vector-symbolic-policy-gradient.jsonld"}}