cd /news/neural-networks/bipropagation-a-decomposition-study-… · home › topics › neural-networks › article
[ARTICLE · art-146061] src=github.com ↗ pub= topic=neural-networks verified=true sentiment=· neutral

Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea

A component-level decomposition study of Dr. Bojan Ploj's bipropagation method finds that per-layer supervision, not gradient locality, is the ingredient that makes greedy layer-wise supervised training effective, with a deeply-supervised control matching the local-loss method at depth 16 on MNIST (0.9684 vs 0.9685). In full-scale runs at 30k train / 10k test with seed 0, local-loss reached 0.9708, 0.9714, 0.9701 and 0.9685 at depths 2, 4, 8 and 16, while vanilla backprop collapsed to 0.1135 at depth 16. On CIFAR-10 (15k train / 10k test, 3 seeds, 12 epochs, no augmentation), the plain end-to-end CNN degraded from 0.643 at depth 6 to 0.557 at depth 9, while greedy local-loss reached 0.626 and deeply-supervised 0.609, indicating a small secondary locality effect on CNNs.

read6 min views1 publishedOct 6, 2026
Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea
Image: Michielbdejong (auto-discovered)

A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's bipropagation idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.

Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand which component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.

The thesis. Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is per-layer supervision, and the approach is competitive and depth-robust on MNIST and CIFAR-10.

To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a single global gradient. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary locality effect appearing on CIFAR-10/CNN.

In short: per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.

Full-scale run, 30k train / 10k test, seed 0:

Depth Vanilla BP Modern BP Anchors (Ploj-style) Local-loss
2 0.9587 0.9690 0.8774 0.9708
4 0.9630 0.9735 0.8743 0.9714
8 0.9586 0.9627 0.8653 0.9701
16 0.1135 (collapse) 0.9513 0.8415 0.9685

The locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that deeply-supervised ≈ local-loss at every depth:

Depth Residual BP Plain BP (ReLU+BN, 30ep) Deeply-supervised (global grad + aux heads) Local-loss
8 0.9674 0.9691 0.9725 0.9701
16 0.9351* 0.9661 0.9684 0.9685

*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.

At depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while restoring the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.

15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):

Depth (blocks) E2E backprop Greedy local-loss Deeply-supervised
3 0.528 ±.016 0.567 ±.003 0.520 ±.022
6 0.643 ±.007 0.649 ±.004 0.577 ±.018
9 0.557 ±.022 ↓ 0.626 ±.005 0.609 ±.013

On CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here local (0.626) edges out deepsup (0.609) at depth 9, so locality contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.

All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.

Method Description
End-to-end (vanilla) backprop Naive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite.
Modern backprop Adam + He init + BatchNorm. The strong baseline.
Greedy local-loss (layer-wise) The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer.
Deeply-supervised control Per-layer auxiliary classifier heads with a single global gradient (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained.
Anchors (Ploj-style) Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout.
Deterministic centroid-init One hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = 0.99·layer_output + 0.01·two_hot(class) , refined with Adam.
.
├── README.md                       # this file
├── PAPER.md                        # the English paper (authoritative findings & numbers)
├── LICENSE                         # MIT
├── requirements.txt
├── .gitignore
├── experiments/
│   ├── cifar_experiment.py         # CIFAR-10 / CNN decomposition (e2e, local, deepsup)
│   └── mnist_mlp_experiment.py     # MNIST / MLP, all methods (self-contained)
└── archive/
    ├── README.md
    └── ...                         # raw development fragments, kept for provenance

The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.

pip install -r requirements.txt
python experiments/cifar_experiment.py        # CIFAR-10 / CNN
python experiments/mnist_mlp_experiment.py    # MNIST / MLP

Upload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.

  • FAST_MODE flag. Each script has aFAST_MODE toggle near the top.True gives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it toFalse for the full benchmark reported in the paper.

  • CIFAR-10 download.cifar_experiment.py downloads CIFAR-10 fromcs.toronto.edu viatf.keras.datasets on first run and caches it to disk; subsequent runs reuse the cache.

  • All reported numbers come from actual evaluation runs.

  • Seeds. Most decisive numbers (full-scale and control runs) are single-seed (seed 0); the multi-seed evidence is currently FAST_MODE / CIFAR only. A fuller protocol (≥10 seeds, 95% CIs, paired Holm-Bonferroni tests) is left for follow-up.

  • Plain, non-residual baselines. Both testbeds compare against plain baselines that degrade with depth for known optimization reasons. A residual/normalized end-to-end baseline would likely close the depth gap, so the depth-robustness claims are relative toplain architectures, not modern residual networks.

  • Iso-compute. Local-loss sees the data roughly 3-6x more often than a single end-to-end run; a clean accuracy-vs-wall-clock and iso-gradient-step accounting is still outstanding.

  • Reconstruction of Ploj's rule. The "anchors" method reconstructs an unpublished multi-class target rule (the originalMNIST.m is auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.

The bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj. This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.

Dr. Ploj's repositories:

Related foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See PAPER.md for the full reference list.

Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:

  • Residual / normalized end-to-end baselines. How does the depth-robustness gap look against a properly modern baseline?
  • More seeds + confidence intervals. ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests on the full-scale and control tables.
  • Harder data. CIFAR-100, Tiny-ImageNet.
  • Other local-learning methods. Forward-Forward, Difference Target Propagation, synthetic gradients, feedback alignment, as additional points of comparison.
  • A more faithful reconstruction of Ploj's multi-class intermediate-target rule (ideally from the originalMNIST.m ).
  • Iso-compute accounting. Accuracy vs. wall-clock and vs. gradient steps.

Open an issue or a pull request.

@misc{korent2026bipropagation,
  author       = {Korent, Maj},
  title        = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},
  year         = {2026},
  howpublished = {\url{https://github.com/korentmaj/bipropagation-study}},
  note         = {Decomposition study building on Bojan Ploj's bipropagation method.}
}

This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.

── more in #neural-networks 4 stories · sorted by recency
── more on @bojan ploj 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/bipropagation-a-deco…] indexed:0 read:6min 2026-10-06 · —