{"slug": "bipropagation-a-decomposition-study-building-on-dr-bojan-ploj-s-idea", "title": "Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea", "summary": "A component-level decomposition study of Dr. Bojan Ploj's bipropagation method finds that per-layer supervision, not gradient locality, is the ingredient that makes greedy layer-wise supervised training effective, with a deeply-supervised control matching the local-loss method at depth 16 on MNIST (0.9684 vs 0.9685). In full-scale runs at 30k train / 10k test with seed 0, local-loss reached 0.9708, 0.9714, 0.9701 and 0.9685 at depths 2, 4, 8 and 16, while vanilla backprop collapsed to 0.1135 at depth 16. On CIFAR-10 (15k train / 10k test, 3 seeds, 12 epochs, no augmentation), the plain end-to-end CNN degraded from 0.643 at depth 6 to 0.557 at depth 9, while greedy local-loss reached 0.626 and deeply-supervised 0.609, indicating a small secondary locality effect on CNNs.", "body_md": "A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's **bipropagation** idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.\n\nBuilding on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand *which* component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.\n\n**The thesis.** Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is **per-layer supervision**, and the approach is competitive and depth-robust on MNIST and CIFAR-10.\n\nTo isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a *single global gradient*. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary *locality* effect appearing on CIFAR-10/CNN.\n\nIn short: **per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.**\n\nFull-scale run, 30k train / 10k test, seed 0:\n\n| Depth | Vanilla BP | Modern BP | Anchors (Ploj-style) | Local-loss | \n|---|---|---|---|---|\n| 2 | 0.9587 | 0.9690 | 0.8774 | **0.9708** | \n| 4 | 0.9630 | 0.9735 | 0.8743 | **0.9714** | \n| 8 | 0.9586 | 0.9627 | 0.8653 | **0.9701** | \n| 16 | 0.1135 (collapse) | 0.9513 | 0.8415 | **0.9685** | \n\nThe locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that **deeply-supervised ≈ local-loss** at every depth:\n\n| Depth | Residual BP | Plain BP (ReLU+BN, 30ep) | Deeply-supervised (global grad + aux heads) | Local-loss | \n|---|---|---|---|---|\n| 8 | 0.9674 | 0.9691 | **0.9725** | 0.9701 | \n| 16 | 0.9351* | 0.9661 | **0.9684** | **0.9685** | \n\n*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.\n\nAt depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while *restoring* the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.\n\n15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):\n\n| Depth (blocks) | E2E backprop | Greedy local-loss | Deeply-supervised | \n|---|---|---|---|\n| 3 | 0.528 ±.016 | **0.567** ±.003 | 0.520 ±.022 | \n| 6 | 0.643 ±.007 | **0.649** ±.004 | 0.577 ±.018 | \n| 9 | 0.557 ±.022 ↓ | **0.626** ±.005 | 0.609 ±.013 | \n\nOn CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here `local` (0.626) edges out `deepsup` (0.609) at depth 9, so *locality* contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.\n\nAll methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.\n\n| Method | Description | \n|---|---|\n| **End-to-end (vanilla) backprop** | Naive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite. | \n| **Modern backprop** | Adam + He init + BatchNorm. The strong baseline. | \n| **Greedy local-loss (layer-wise)** | The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer. | \n| **Deeply-supervised control** | Per-layer auxiliary classifier heads with a *single global gradient* (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained. | \n| **Anchors (Ploj-style)** | Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout. | \n| **Deterministic centroid-init** | One hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = `0.99·layer_output + 0.01·two_hot(class)` , refined with Adam. | \n\n```\n.\n├── README.md                       # this file\n├── PAPER.md                        # the English paper (authoritative findings & numbers)\n├── LICENSE                         # MIT\n├── requirements.txt\n├── .gitignore\n├── experiments/\n│   ├── cifar_experiment.py         # CIFAR-10 / CNN decomposition (e2e, local, deepsup)\n│   └── mnist_mlp_experiment.py     # MNIST / MLP, all methods (self-contained)\n└── archive/\n    ├── README.md\n    └── ...                         # raw development fragments, kept for provenance\n```\n\nThe experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.\n\n```\npip install -r requirements.txt\npython experiments/cifar_experiment.py        # CIFAR-10 / CNN\npython experiments/mnist_mlp_experiment.py    # MNIST / MLP\n```\n\nUpload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.\n\n- **`FAST_MODE` flag.** Each script has a`FAST_MODE` toggle near the top.`True` gives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it to`False` for the full benchmark reported in the paper.\n- **CIFAR-10 download.**`cifar_experiment.py` downloads CIFAR-10 from`cs.toronto.edu` via`tf.keras.datasets` on first run and caches it to disk; subsequent runs reuse the cache.\n- All reported numbers come from actual evaluation runs.\n\n- **Seeds.** Most decisive numbers (full-scale and control runs) are single-seed (seed 0); the multi-seed evidence is currently FAST_MODE / CIFAR only. A fuller protocol (≥10 seeds, 95% CIs, paired Holm-Bonferroni tests) is left for follow-up.\n- **Plain, non-residual baselines.** Both testbeds compare against plain baselines that degrade with depth for known optimization reasons. A residual/normalized end-to-end baseline would likely close the depth gap, so the depth-robustness claims are relative to*plain* architectures, not modern residual networks.\n- **Iso-compute.** Local-loss sees the data roughly 3-6x more often than a single end-to-end run; a clean accuracy-vs-wall-clock and iso-gradient-step accounting is still outstanding.\n- **Reconstruction of Ploj's rule.** The \"anchors\" method reconstructs an unpublished multi-class target rule (the original`MNIST.m` is auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.\n\nThe **bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj.** This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.\n\nDr. Ploj's repositories:\n\nRelated foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See `PAPER.md` for the full reference list.\n\nContributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:\n\n- **Residual / normalized end-to-end baselines.** How does the depth-robustness gap look against a properly modern baseline?\n- **More seeds + confidence intervals.** ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests on the full-scale and control tables.\n- **Harder data.** CIFAR-100, Tiny-ImageNet.\n- **Other local-learning methods.** Forward-Forward, Difference Target Propagation, synthetic gradients, feedback alignment, as additional points of comparison.\n- **A more faithful reconstruction** of Ploj's multi-class intermediate-target rule (ideally from the original`MNIST.m` ).\n- **Iso-compute accounting.** Accuracy vs. wall-clock and vs. gradient steps.\n\nOpen an issue or a pull request.\n\n```\n@misc{korent2026bipropagation,\n  author       = {Korent, Maj},\n  title        = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},\n  year         = {2026},\n  howpublished = {\\url{https://github.com/korentmaj/bipropagation-study}},\n  note         = {Decomposition study building on Bojan Ploj's bipropagation method.}\n}\n```\n\n*This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.*", "url": "https://wpnews.pro/news/bipropagation-a-decomposition-study-building-on-dr-bojan-ploj-s-idea", "canonical_source": "https://github.com/korentmaj/bipropagation-study", "published_at": "2026-10-06 13:10:50+00:00", "updated_at": "2026-10-06 13:21:08.567351+00:00", "lang": "en", "topics": ["neural-networks", "machine-learning", "ai-research"], "entities": ["Bojan Ploj", "bipropagation", "MNIST", "CIFAR-10", "Belilovsky", "Nøkland", "Lee"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/bipropagation-a-decomposition-study-building-on-dr-bojan-ploj-s-idea", "markdown": "https://wpnews.pro/news/bipropagation-a-decomposition-study-building-on-dr-bojan-ploj-s-idea.md", "text": "https://wpnews.pro/news/bipropagation-a-decomposition-study-building-on-dr-bojan-ploj-s-idea.txt", "jsonld": "https://wpnews.pro/news/bipropagation-a-decomposition-study-building-on-dr-bojan-ploj-s-idea.jsonld"}}