Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea A component-level decomposition study of Dr. Bojan Ploj's bipropagation method finds that per-layer supervision, not gradient locality, is the ingredient that makes greedy layer-wise supervised training effective, with a deeply-supervised control matching the local-loss method at depth 16 on MNIST (0.9684 vs 0.9685). In full-scale runs at 30k train / 10k test with seed 0, local-loss reached 0.9708, 0.9714, 0.9701 and 0.9685 at depths 2, 4, 8 and 16, while vanilla backprop collapsed to 0.1135 at depth 16. On CIFAR-10 (15k train / 10k test, 3 seeds, 12 epochs, no augmentation), the plain end-to-end CNN degraded from 0.643 at depth 6 to 0.557 at depth 9, while greedy local-loss reached 0.626 and deeply-supervised 0.609, indicating a small secondary locality effect on CNNs. A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's bipropagation idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets. Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand which component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10. The thesis. Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is per-layer supervision , and the approach is competitive and depth-robust on MNIST and CIFAR-10. To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a single global gradient . This control matches the local-loss method at every depth depth-16: 0.9684 vs 0.9685 , which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary locality effect appearing on CIFAR-10/CNN. In short: per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper. Full-scale run, 30k train / 10k test, seed 0: | Depth | Vanilla BP | Modern BP | Anchors Ploj-style | Local-loss | |---|---|---|---|---| | 2 | 0.9587 | 0.9690 | 0.8774 | 0.9708 | | 4 | 0.9630 | 0.9735 | 0.8743 | 0.9714 | | 8 | 0.9586 | 0.9627 | 0.8653 | 0.9701 | | 16 | 0.1135 collapse | 0.9513 | 0.8415 | 0.9685 | The locality-isolation control 30k MNIST, seed 0, 30 epochs shows that deeply-supervised ≈ local-loss at every depth: | Depth | Residual BP | Plain BP ReLU+BN, 30ep | Deeply-supervised global grad + aux heads | Local-loss | |---|---|---|---|---| | 8 | 0.9674 | 0.9691 | 0.9725 | 0.9701 | | 16 | 0.9351 | 0.9661 | 0.9684 | 0.9685 | The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline. At depth 16, deeply-supervised 0.9684 ≈ local-loss 0.9685 : keeping per-layer supervision while restoring the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient. 15k train / 10k test, 3 seeds, 12 epochs, no augmentation held identical across methods : | Depth blocks | E2E backprop | Greedy local-loss | Deeply-supervised | |---|---|---|---| | 3 | 0.528 ±.016 | 0.567 ±.003 | 0.520 ±.022 | | 6 | 0.643 ±.007 | 0.649 ±.004 | 0.577 ±.018 | | 9 | 0.557 ±.022 ↓ | 0.626 ±.005 | 0.609 ±.013 | On CIFAR the plain non-residual end-to-end CNN degrades at depth 9 0.643 to 0.557 . Both per-layer-supervised methods are more depth-robust, and here local 0.626 edges out deepsup 0.609 at depth 9, so locality contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision. All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds. | Method | Description | |---|---| | End-to-end vanilla backprop | Naive init, saturating tanh activation, plain SGD. This is the regime where vanishing gradients bite. | | Modern backprop | Adam + He init + BatchNorm. The strong baseline. | | Greedy local-loss layer-wise | The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy à la Belilovsky 2019 / Nøkland 2019 ; the head is discarded before the next layer. | | Deeply-supervised control | Per-layer auxiliary classifier heads with a single global gradient Lee 2015 . The locality-isolation control: same per-layer supervision, but global backprop is retained. | | Anchors Ploj-style | Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout. | | Deterministic centroid-init | One hidden layer constructed analytically from class-centroid geometry sparse ±1 units over the 3 most-discriminative features , targets = 0.99·layer output + 0.01·two hot class , refined with Adam. | . ├── README.md this file ├── PAPER.md the English paper authoritative findings & numbers ├── LICENSE MIT ├── requirements.txt ├── .gitignore ├── experiments/ │ ├── cifar experiment.py CIFAR-10 / CNN decomposition e2e, local, deepsup │ └── mnist mlp experiment.py MNIST / MLP, all methods self-contained └── archive/ ├── README.md └── ... raw development fragments, kept for provenance The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally. pip install -r requirements.txt python experiments/cifar experiment.py CIFAR-10 / CNN python experiments/mnist mlp experiment.py MNIST / MLP Upload a script or paste it into a cell and run. A GPU runtime e.g. T4 is recommended for the full configs. - FAST MODE flag. Each script has a FAST MODE toggle near the top. True gives a small, fast indicative smoke run subset of data, few epochs, 2 to 3 seeds ; set it to False for the full benchmark reported in the paper. - CIFAR-10 download. cifar experiment.py downloads CIFAR-10 from cs.toronto.edu via tf.keras.datasets on first run and caches it to disk; subsequent runs reuse the cache. - All reported numbers come from actual evaluation runs. - Seeds. Most decisive numbers full-scale and control runs are single-seed seed 0 ; the multi-seed evidence is currently FAST MODE / CIFAR only. A fuller protocol ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests is left for follow-up. - Plain, non-residual baselines. Both testbeds compare against plain baselines that degrade with depth for known optimization reasons. A residual/normalized end-to-end baseline would likely close the depth gap, so the depth-robustness claims are relative to plain architectures, not modern residual networks. - Iso-compute. Local-loss sees the data roughly 3-6x more often than a single end-to-end run; a clean accuracy-vs-wall-clock and iso-gradient-step accounting is still outstanding. - Reconstruction of Ploj's rule. The "anchors" method reconstructs an unpublished multi-class target rule the original MNIST.m is auth-walled on ResearchGate . A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion. The bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj. This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible. Dr. Ploj's repositories: Related foundational work this study builds on includes Deeply-Supervised Nets Lee et al. 2015 , greedy layer-wise learning at scale Belilovsky et al. 2019 , local error signals Nøkland & Eidnes 2019 , and Difference Target Propagation Lee et al. 2015 . See PAPER.md for the full reference list. Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions: - Residual / normalized end-to-end baselines. How does the depth-robustness gap look against a properly modern baseline? - More seeds + confidence intervals. ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests on the full-scale and control tables. - Harder data. CIFAR-100, Tiny-ImageNet. - Other local-learning methods. Forward-Forward, Difference Target Propagation, synthetic gradients, feedback alignment, as additional points of comparison. - A more faithful reconstruction of Ploj's multi-class intermediate-target rule ideally from the original MNIST.m . - Iso-compute accounting. Accuracy vs. wall-clock and vs. gradient steps. Open an issue or a pull request. @misc{korent2026bipropagation, author = {Korent, Maj}, title = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea}, year = {2026}, howpublished = {\url{https://github.com/korentmaj/bipropagation-study}}, note = {Decomposition study building on Bojan Ploj's bipropagation method.} } This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.