{"slug": "program-learning-with-verifiable-rewards-symbolic-backpropagation", "title": "Program Learning with Verifiable Rewards: Symbolic Backpropagation", "summary": "Researchers introduced PLVR (Program Learning with Verifiable Rewards), a post-training method that learns explicit programs from input-output examples via symbolic backpropagation, placing reasoning outside a model's weights. On LiveCodeBench v6 and Tau2Bench, 30B base models using PLVR outperformed RL at matched budget by 27.8 points on average and frontier models an order of magnitude larger by 13.6 points. The authors released the symbolic backpropagation library and a conformance checker.", "body_md": "# Computer Science > Artificial Intelligence\n\n[Submitted on 28 Aug 2026]\n\n# Title:Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs\n\n[View PDF](/pdf/2608.28421)\n\n[HTML (experimental)](https://arxiv.org/html/2608.28421v1)\n\nAbstract:Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by step and cannot be moved to another model. We argue that for tasks whose intermediate steps admit verification, reasoning is better placed outside the base models weights as an explicit program composed from deterministic and neural primitives. We introduce PLVR (Program Learning with Verifiable Rewards): a post training method that learns such programs directly from input-output examples. Its mechanism is symbolic backpropagation: each program layer carries a typed ontology a loss is computed at the output against ground truth and required input ontologies are propagated backward by type inference over primitive signatures: an analogue of the chain rule in which credit assignment is a derivation rather than an estimate. Where RLVR verifies a terminal outcome, PLVRs reward is a per step contract verdict dense over program structure. On LiveCodeBench v6 and Tau2Bench, 30B base models with PLVR outperform RL at matched budget by 27.8 points on average and frontier models an order of magnitude larger by 13.6 points. A single primitive library serves two benchmarks, so the marginal cost of a new task is 100 examples of program search and no new finetuning data. Replacing the loss guided search with uniform sampling over the same type admissible space at equal budget collapses the median program from 65.6 to 17.5, identifying the backward pass rather than the type system as the source of the advantage. We release the symbolic backpropagation library and a conformance checker so the method can be applied to primitive libraries other than our own.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/program-learning-with-verifiable-rewards-symbolic-backpropagation", "canonical_source": "https://arxiv.org/abs/2608.28421", "published_at": "2026-09-01 04:07:07+00:00", "updated_at": "2026-09-01 04:22:55.647725+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["PLVR", "LiveCodeBench v6", "Tau2Bench"], "alternates": {"html": "https://wpnews.pro/news/program-learning-with-verifiable-rewards-symbolic-backpropagation", "markdown": "https://wpnews.pro/news/program-learning-with-verifiable-rewards-symbolic-backpropagation.md", "text": "https://wpnews.pro/news/program-learning-with-verifiable-rewards-symbolic-backpropagation.txt", "jsonld": "https://wpnews.pro/news/program-learning-with-verifiable-rewards-symbolic-backpropagation.jsonld"}}