cd /news/large-language-models/bp-llm-belief-propagation-for-binary… · home › topics › large-language-models › article
[ARTICLE · art-147243] src=aclanthology.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

BP-LLM: Belief Propagation for Binary Feedback in Large Language Model Alignment

Jessica E. Liang published BP-LLM, a probabilistic framework that models binary preference feedback as noisy observations of latent reward margins under a policy-induced Gaussian prior, in Transactions of the Association for Computational Linguistics volume 14, pages 1413–1430. BP-LLM uses the Jaakkola–Jordan variational bound to derive closed-form Gaussian message updates and recovers Binary Classifier Optimization and Direct Preference Optimization as special cases. In inference-only tests with frozen weights, BP-LLM improved label-free test-time win rate over BCO and DPO across UltraFeedback, Capybara, and HelpSteer2 for open-weight Llama and Qwen models, and with parameter-efficient LoRA updates it outperformed Cal-DPO on win rate.

read2 min views1 publishedOct 7, 2026
BP-LLM: Belief Propagation for Binary Feedback in Large Language Model Alignment
Image: Aclanthology (auto-discovered)
Abstract

Preference-based alignment methods such as Direct Preference Optimization (DPO) and Binary Classifier Optimization (BCO) offer efficient alternatives to reinforcement learning from human feedback (RLHF). However, they often treat comparisons as independent labels and do not explicitly model uncertainty, which can lead to miscalibration and reduced robustness under noisy or heterogeneous feedback. We introduce Belief Propagation for Large Language Model Alignment (BP–LLM), a probabilistic framework that views binary feedback as noisy observations of latent reward margins under a policy-induced Gaussian prior. Using the Jaakkola–Jordan variational bound, BP–LLM yields closed-form Gaussian message updates and performs stable belief propagation between a classifier-side inference module and the policy. Exchanging extrinsic messages enables both modules to refine beliefs without double counting and recovers BCO and DPO as special cases. We evaluate BP–LLM in two regimes. In an inference-only setting with frozen LLM weights, BP–LLM consistently improves label-free test-time win rate over BCO and DPO across UltraFeedback, Capybara, and HelpSteer2 for open-weight Llama and Qwen models. In a training-time setting with parameter-efficient LoRA updates, BP–LLM also outperforms Cal-DPO with higher win rates. Overall, BP–LLM is most beneficial under noisy/heterogeneous (or unary) feedback, where posterior refinement and extrinsic shaping provide more reliable signals than hard labels, while remaining lightweight and scalable.1

- Anthology ID:
- 2026.tacl-1.64
- Volume:
- [Transactions of the Association for Computational Linguistics, Volume 14](https://aclanthology.org/volumes/2026.tacl-1/)
- Month:
- Year:
  • 2026
  • Address:
  • Cambridge, MA
- Venue:
- [TACL](https://aclanthology.org/venues/tacl/)
- SIG:
- Publisher:
  • MIT Press
- Note:
- Pages:
  • 1413–1430
- Language:
- URL:
- [https://aclanthology.org/2026.tacl-1.64/](https://aclanthology.org/2026.tacl-1.64/)
- DOI:
- [10.1162/tacl.a.734](https://doi.org/10.1162/tacl.a.734)
- Cite (ACL):
- Cite (Informal):
- [BP-LLM: Belief Propagation for Binary Feedback in Large Language Model Alignment](https://aclanthology.org/2026.tacl-1.64/) (Liang, TACL 2026)
- PDF:
- [https://aclanthology.org/2026.tacl-1.64.pdf](https://aclanthology.org/2026.tacl-1.64.pdf)
── more in #large-language-models 4 stories · sorted by recency
── more on @jessica e. liang 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/bp-llm-belief-propag…] indexed:0 read:2min 2026-10-07 · —