# Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea

> Source: <https://github.com/korentmaj/bipropagation-study>
> Published: 2026-10-06 13:10:50+00:00

A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's **bipropagation** idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.

Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand *which* component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.

**The thesis.** Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is **per-layer supervision**, and the approach is competitive and depth-robust on MNIST and CIFAR-10.

To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a *single global gradient*. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary *locality* effect appearing on CIFAR-10/CNN.

In short: **per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.**

Full-scale run, 30k train / 10k test, seed 0:

| Depth | Vanilla BP | Modern BP | Anchors (Ploj-style) | Local-loss | 
|---|---|---|---|---|
| 2 | 0.9587 | 0.9690 | 0.8774 | **0.9708** | 
| 4 | 0.9630 | 0.9735 | 0.8743 | **0.9714** | 
| 8 | 0.9586 | 0.9627 | 0.8653 | **0.9701** | 
| 16 | 0.1135 (collapse) | 0.9513 | 0.8415 | **0.9685** | 

The locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that **deeply-supervised ≈ local-loss** at every depth:

| Depth | Residual BP | Plain BP (ReLU+BN, 30ep) | Deeply-supervised (global grad + aux heads) | Local-loss | 
|---|---|---|---|---|
| 8 | 0.9674 | 0.9691 | **0.9725** | 0.9701 | 
| 16 | 0.9351* | 0.9661 | **0.9684** | **0.9685** | 

*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.

At depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while *restoring* the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.

15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):

| Depth (blocks) | E2E backprop | Greedy local-loss | Deeply-supervised | 
|---|---|---|---|
| 3 | 0.528 ±.016 | **0.567** ±.003 | 0.520 ±.022 | 
| 6 | 0.643 ±.007 | **0.649** ±.004 | 0.577 ±.018 | 
| 9 | 0.557 ±.022 ↓ | **0.626** ±.005 | 0.609 ±.013 | 

On CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here `local` (0.626) edges out `deepsup` (0.609) at depth 9, so *locality* contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.

All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.

| Method | Description | 
|---|---|
| **End-to-end (vanilla) backprop** | Naive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite. | 
| **Modern backprop** | Adam + He init + BatchNorm. The strong baseline. | 
| **Greedy local-loss (layer-wise)** | The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer. | 
| **Deeply-supervised control** | Per-layer auxiliary classifier heads with a *single global gradient* (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained. | 
| **Anchors (Ploj-style)** | Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout. | 
| **Deterministic centroid-init** | One hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = `0.99·layer_output + 0.01·two_hot(class)` , refined with Adam. | 

```
.
├── README.md                       # this file
├── PAPER.md                        # the English paper (authoritative findings & numbers)
├── LICENSE                         # MIT
├── requirements.txt
├── .gitignore
├── experiments/
│   ├── cifar_experiment.py         # CIFAR-10 / CNN decomposition (e2e, local, deepsup)
│   └── mnist_mlp_experiment.py     # MNIST / MLP, all methods (self-contained)
└── archive/
    ├── README.md
    └── ...                         # raw development fragments, kept for provenance
```

The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.

```
pip install -r requirements.txt
python experiments/cifar_experiment.py        # CIFAR-10 / CNN
python experiments/mnist_mlp_experiment.py    # MNIST / MLP
```

Upload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.

- **`FAST_MODE` flag.** Each script has a`FAST_MODE` toggle near the top.`True` gives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it to`False` for the full benchmark reported in the paper.
- **CIFAR-10 download.**`cifar_experiment.py` downloads CIFAR-10 from`cs.toronto.edu` via`tf.keras.datasets` on first run and caches it to disk; subsequent runs reuse the cache.
- All reported numbers come from actual evaluation runs.

- **Seeds.** Most decisive numbers (full-scale and control runs) are single-seed (seed 0); the multi-seed evidence is currently FAST_MODE / CIFAR only. A fuller protocol (≥10 seeds, 95% CIs, paired Holm-Bonferroni tests) is left for follow-up.
- **Plain, non-residual baselines.** Both testbeds compare against plain baselines that degrade with depth for known optimization reasons. A residual/normalized end-to-end baseline would likely close the depth gap, so the depth-robustness claims are relative to*plain* architectures, not modern residual networks.
- **Iso-compute.** Local-loss sees the data roughly 3-6x more often than a single end-to-end run; a clean accuracy-vs-wall-clock and iso-gradient-step accounting is still outstanding.
- **Reconstruction of Ploj's rule.** The "anchors" method reconstructs an unpublished multi-class target rule (the original`MNIST.m` is auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.

The **bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj.** This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.

Dr. Ploj's repositories:

Related foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See `PAPER.md` for the full reference list.

Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:

- **Residual / normalized end-to-end baselines.** How does the depth-robustness gap look against a properly modern baseline?
- **More seeds + confidence intervals.** ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests on the full-scale and control tables.
- **Harder data.** CIFAR-100, Tiny-ImageNet.
- **Other local-learning methods.** Forward-Forward, Difference Target Propagation, synthetic gradients, feedback alignment, as additional points of comparison.
- **A more faithful reconstruction** of Ploj's multi-class intermediate-target rule (ideally from the original`MNIST.m` ).
- **Iso-compute accounting.** Accuracy vs. wall-clock and vs. gradient steps.

Open an issue or a pull request.

```
@misc{korent2026bipropagation,
  author       = {Korent, Maj},
  title        = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},
  year         = {2026},
  howpublished = {\url{https://github.com/korentmaj/bipropagation-study}},
  note         = {Decomposition study building on Bojan Ploj's bipropagation method.}
}
```

*This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.*
