cd /news/machine-learning/reinforcement-learning-and-rule-base… · home topics machine-learning article
[ARTICLE · art-119766] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Reinforcement Learning and Rule-Based Peer-to-Peer Pricing in Residential PV-BES Communities

A new arXiv paper (2609.01680v1) comparing rule-based and reinforcement-learning (RL) pricing for peer-to-peer electricity trading in residential photovoltaic communities finds that rule-based benchmarks outperform the best RL policy in the base PV-only configuration, while adding battery storage boosts community savings under the best RL policy from EUR 734.23 to EUR 978.52. The study, which uses a Deep Q-Network and evaluates multiplier-based and learnable supply-demand-ratio (SDR) shaped pricing, concludes that rule-based pricing remains highly competitive and that storage significantly improves learning-based outcomes, though benefit distribution remains heterogeneous across households.

read1 min views1 publishedSep 3, 2026

arXiv:2609.01680v1 Announce Type: new Abstract: This paper compares rule-based and learning-based pricing mechanisms for peer-to-peer (P2P) electricity trading in residential photovoltaic communities. The rule-based benchmarks comprise bill-sharing as an ex post allocation mechanism, the mid-market rate, and supply-demand-ratio pricing. The reinforcement-learning (RL) formulation is implemented through a Deep Q-Network and evaluated under multiplier-based and learnable SDR-shaped pricing, with a fixed-parameter SDR variant as a non-learning control. Performance is assessed through community savings together with complementary financial and operational indicators. In the base PV-only configuration, the rule-based benchmarks outperform the best RL policy. With battery energy storage, evaluated for the RL policies only, community savings under the best RL policy increase from EUR 734.23 to EUR 978.52. Across the learning-based modes and in both configurations, SDR-shaped pricing outperforms the multiplier-based parameterization considered. The results indicate that rule-based pricing remains highly competitive wherever the two families are compared directly, and that storage substantially improves the learning-based outcomes under this accounting, while the distribution of benefits remains heterogeneous across households.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reinforcement-learni…] indexed:0 read:1min 2026-09-03 ·