A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's bipropagation idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.
Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand which component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.
The thesis. Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is per-layer supervision, and the approach is competitive and depth-robust on MNIST and CIFAR-10.
To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a single global gradient. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary locality effect appearing on CIFAR-10/CNN.
In short: per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.
Full-scale run, 30k train / 10k test, seed 0:
| Depth | Vanilla BP | Modern BP | Anchors (Ploj-style) | Local-loss |
|---|---|---|---|---|
| 2 | 0.9587 | 0.9690 | 0.8774 | 0.9708 |
| 4 | 0.9630 | 0.9735 | 0.8743 | 0.9714 |
| 8 | 0.9586 | 0.9627 | 0.8653 | 0.9701 |
| 16 | 0.1135 (collapse) | 0.9513 | 0.8415 | 0.9685 |
The locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that deeply-supervised ≈ local-loss at every depth:
| Depth | Residual BP | Plain BP (ReLU+BN, 30ep) | Deeply-supervised (global grad + aux heads) | Local-loss |
|---|---|---|---|---|
| 8 | 0.9674 | 0.9691 | 0.9725 | 0.9701 |
| 16 | 0.9351* | 0.9661 | 0.9684 | 0.9685 |
*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.
At depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while restoring the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.
15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):
| Depth (blocks) | E2E backprop | Greedy local-loss | Deeply-supervised |
|---|---|---|---|
| 3 | 0.528 ±.016 | 0.567 ±.003 | 0.520 ±.022 |
| 6 | 0.643 ±.007 | 0.649 ±.004 | 0.577 ±.018 |
| 9 | 0.557 ±.022 ↓ | 0.626 ±.005 | 0.609 ±.013 |
On CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here local (0.626) edges out deepsup (0.609) at depth 9, so locality contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.
All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.
| Method | Description |
|---|---|
| End-to-end (vanilla) backprop | Naive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite. |
| Modern backprop | Adam + He init + BatchNorm. The strong baseline. |
| Greedy local-loss (layer-wise) | The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer. |
| Deeply-supervised control | Per-layer auxiliary classifier heads with a single global gradient (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained. |
| Anchors (Ploj-style) | Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout. |
| Deterministic centroid-init | One hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = 0.99·layer_output + 0.01·two_hot(class) , refined with Adam. |
.
├── README.md # this file
├── PAPER.md # the English paper (authoritative findings & numbers)
├── LICENSE # MIT
├── requirements.txt
├── .gitignore
├── experiments/
│ ├── cifar_experiment.py # CIFAR-10 / CNN decomposition (e2e, local, deepsup)
│ └── mnist_mlp_experiment.py # MNIST / MLP, all methods (self-contained)
└── archive/
├── README.md
└── ... # raw development fragments, kept for provenance
The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.
pip install -r requirements.txt
python experiments/cifar_experiment.py # CIFAR-10 / CNN
python experiments/mnist_mlp_experiment.py # MNIST / MLP
Upload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.
-
FAST_MODEflag. Each script has aFAST_MODEtoggle near the top.Truegives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it toFalsefor the full benchmark reported in the paper. -
CIFAR-10 download.
cifar_experiment.pydownloads CIFAR-10 fromcs.toronto.eduviatf.keras.datasetson first run and caches it to disk; subsequent runs reuse the cache. -
All reported numbers come from actual evaluation runs.
-
Seeds. Most decisive numbers (full-scale and control runs) are single-seed (seed 0); the multi-seed evidence is currently FAST_MODE / CIFAR only. A fuller protocol (≥10 seeds, 95% CIs, paired Holm-Bonferroni tests) is left for follow-up.
-
Plain, non-residual baselines. Both testbeds compare against plain baselines that degrade with depth for known optimization reasons. A residual/normalized end-to-end baseline would likely close the depth gap, so the depth-robustness claims are relative toplain architectures, not modern residual networks.
-
Iso-compute. Local-loss sees the data roughly 3-6x more often than a single end-to-end run; a clean accuracy-vs-wall-clock and iso-gradient-step accounting is still outstanding.
-
Reconstruction of Ploj's rule. The "anchors" method reconstructs an unpublished multi-class target rule (the original
MNIST.mis auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.
The bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj. This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.
Dr. Ploj's repositories:
Related foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See PAPER.md for the full reference list.
Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:
- Residual / normalized end-to-end baselines. How does the depth-robustness gap look against a properly modern baseline?
- More seeds + confidence intervals. ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests on the full-scale and control tables.
- Harder data. CIFAR-100, Tiny-ImageNet.
- Other local-learning methods. Forward-Forward, Difference Target Propagation, synthetic gradients, feedback alignment, as additional points of comparison.
- A more faithful reconstruction of Ploj's multi-class intermediate-target rule (ideally from the original
MNIST.m). - Iso-compute accounting. Accuracy vs. wall-clock and vs. gradient steps.
Open an issue or a pull request.
@misc{korent2026bipropagation,
author = {Korent, Maj},
title = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},
year = {2026},
howpublished = {\url{https://github.com/korentmaj/bipropagation-study}},
note = {Decomposition study building on Bojan Ploj's bipropagation method.}
}
This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.