{"slug": "show-hn-dynamictune-closed-form-trajectory-weight-surgery-from-4b-into-0-8b", "title": "Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B", "summary": "DynamicTune, a Show HN project, reports that closed-form weight surgery can transfer trajectory dynamics from a 4B-parameter Qwen3.5 (d=2560) into a 0.8B student (d=1024) and from GPT-2 XL (d=1600) into GPT-2 small (d=768) without backpropagation or training runs. Editing all 24 student layers raised negative log-likelihood by 64.78%, but restricting the surgery to 4 anchor blocks (layers 0, 7, 15, 23) cut multi-domain held-out NLL by 10.8% and improved HellaSwag by 0.50% across 400 tasks on native Vulkan llama.cpp. The project attributes the failure of full-stack edits to high spectral entropy (>0.90) in intermediate layers 1-22 and provides reproducible benchmark logs in its benchmarks directory.", "body_md": "Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths.\n\nTested across radically different model families: modern hybrid `Qwen3.5` (4B with `d=2560` -> 0.8B with `d=1024`) and notoriously fragile `GPT-2` (XL with `d=1600` -> small with `d=768`).\n\nThe prevailing consensus in deep learning is that transferring capabilities from a larger teacher model to a smaller student demands billions or trillions of tokens, massive synthetic dataset pipelines, and weeks of GPU cluster compute running token-level cross-entropy or KL divergence minimization.\n\n**DynamicTune challenges this dogma:**\n\n1. A transformer stack is fundamentally a discrete dynamical system over depth: $h_{l+1} = h_l + f_l(h_l)$ .\n2. A capable teacher traces an informational velocity field through representation space.\n3. By aligning these trajectories through a local orthogonal Procrustes atlas and solving for closed-form weight updates in the student's MLP blocks, we can physically transfer teacher trajectory dynamics into the student without backpropagation or training runs.\n4. \n**Cross-architecture stability:** We tested this on both modern`Qwen3.5` (SwiGLU, hybrid linear/sliding attention) and the notoriously fragile`GPT-2 small` (where Conv1D layers famously collapse or degrade into gibberish at the slightest weight disturbance). In both architectures, baseline language modeling integrity is preserved with well-behaved, bounded degradation margins, while target domain accuracy improves.\n5. \n**The spectral entropy discovery:** Editing all 24 student layers destroys the model (+64.78% NLL) because intermediate layers (1-22) are high-entropy polysemantic knots (>0.90 spectral entropy). Restricting the surgery to**4 anchor blocks** (layers 0, 7, 15, and 23) avoids destructive interference, cuts multi-domain held-out NLL by**-10.8%** , and boosts**HellaSwag (+0.50% across 400 tasks)** on native Vulkan`llama.cpp` .\n\nRaw reproducible benchmark logs: [`benchmarks/`](https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks).\n\nDynamicTune bridges established machine learning theory and mechanistic interpretability into an empirical weight surgery engine:\n\nResidual connections allow layers to be viewed as Euler discretization steps of an underlying continuous ordinary differential equation \n\n- [Chen et al., 2018: Neural Ordinary Differential Equations (arXiv:1806.07366)](https://arxiv.org/abs/1806.07366)\n- [Lu et al., 2017: Beyond Finite Layer Neural Networks: Bridging Deep Architectures and Numerical Differential Equations (arXiv:1710.10121)](https://arxiv.org/abs/1710.10121)\n- [Sander et al., 2022: Residual Neural Networks as Approximations of Ordinary Differential Equations (arXiv:2202.10512)](https://arxiv.org/abs/2202.10512)\n\nIn DynamicTune, we do not view weights as static feature matrices. We treat the step \n\nConcepts in large language models are represented as linear directions in representation space, and different models often learn linearly or orthogonally equivalent geometries up to rotation and scaling.\n\n- [Park et al., 2023: The Linear Representation Hypothesis and the Geometry of Large Language Models (arXiv:2311.03658)](https://arxiv.org/abs/2311.03658)\n- [Kornblith et al., 2019: Similarity of Neural Network Representations Revisited (arXiv:1905.00414)](https://arxiv.org/abs/1905.00414)\n- [Ding et al., 2021: Grounding Representation Similarity with Statistical Mechanics (arXiv:2106.11561)](https://arxiv.org/abs/2106.11561)\n\nBecause teacher and student models have different hidden dimensions (e.g. 2560 vs 1024), a single global orthogonal matrix cannot capture non-linear curvature across different semantic clusters. DynamicTune builds a piecewise local Procrustes atlas (`ManifoldChartAtlas` in `faytuna_flow/manifold_charts.py`): we cluster hidden states with K-Means into \n\nWhy did past attempts at layer-wise weight transfer fail? Anthropic's research into mechanistic interpretability showed that neural networks pack more features than they have dimensions via superposition, creating polysemantic neurons that activate on multiple unrelated concepts.\n\n- [Elhage et al., 2022: Toy Models of Superposition (arXiv:2209.10652)](https://arxiv.org/abs/2209.10652)\n- [Bricken et al., 2023: Towards Monosemanticity: Decomposing Language Models With Dictionary Learning](https://transformer-circuits.pub/2023/monosemantic-features/index.html)\n\nWhen a student model has only 1024 dimensions, intermediate layers (layers 1 to 22) are forced to operate in dense superposition. In DynamicTune, we compute the singular value distribution of the flow residual and calculate its normalized Shannon spectral entropy `faytuna_flow/knots.py`).\n\n- \n**Layer 0** exhibits low entropy ($H = 0.7138$ ): clean, coherent semantic grounding.\n- \n**Layers 1-22** exhibit high entropy ($H \\in [0.8957, 0.9669]$ ): chaotic superposition knots. Forcing a linear weight update here creates catastrophic interference and ruins the model.\n- \n**Layer 23** ($H = 0.9310$ ): pre-unembed boundary where features unpack toward vocabulary logits.\n\nBy discovering this entropy barrier, we learned that weight surgery must respect superposition boundaries: edit sparse anchor points, bypass the knots.\n\nInstead of gradient descent, direct weight updates can be computed as closed-form linear projections that satisfy key-value associations.\n\n- [Meng et al., 2022: Locating and Editing Factual Associations in GPT (ROME, arXiv:2202.05262)](https://arxiv.org/abs/2202.05262)\n- [Meng et al., 2022: Mass-Editing Memory in a Transformer (MEMIT, arXiv:2210.07229)](https://arxiv.org/abs/2210.07229)\n\nDynamicTune extends this concept from individual fact-editing to depth-wise dynamical flow transport: we pull the projected trajectory deltas back through the SwiGLU MLP blocks via a damped Tikhonov pseudoinverse and rank-constrained SVD projections with explicit spectral trust-region bounds.\n\nRunning `scripts/scan_24_layers_autogate.py` across all layers of `Qwen3.5-0.8B` mapped to `Qwen3.5-4B` reveals why full-model transfer fails:\n\n| Student Layer | Mapped Teacher Layer | Multi-Chart Atlas Spectral Entropy | Diagnosis | \n|---|---|---|---|\n| **Layer 0** | **Layer 0** | **0.7138** | **Coherent semantic anchor (safe for surgery)** | \n| Layer 1 | Layer 1 | 0.9669 | Polysemantic knot (skip) | \n| Layer 2 | Layer 3 | 0.9419 | Polysemantic knot (skip) | \n| Layer 3 | Layer 4 | 0.9032 | Polysemantic knot (skip) | \n| Layer 4 | Layer 5 | 0.9572 | Polysemantic knot (skip) | \n| Layer 5 | Layer 7 | 0.9250 | Polysemantic knot (skip) | \n| Layer 6 | Layer 8 | 0.9388 | Polysemantic knot (skip) | \n| Layer 7 | Layer 9 | 0.9384 | Intermediate bridge anchor (damped) | \n| Layer 8 | Layer 11 | 0.9387 | Polysemantic knot (skip) | \n| Layer 9 | Layer 12 | 0.9531 | Polysemantic knot (skip) | \n| Layer 10 | Layer 13 | 0.9182 | Polysemantic knot (skip) | \n| Layer 11 | Layer 15 | 0.9482 | Polysemantic knot (skip) | \n| Layer 12 | Layer 16 | 0.9504 | Polysemantic knot (skip) | \n| Layer 13 | Layer 17 | 0.9183 | Polysemantic knot (skip) | \n| Layer 14 | Layer 19 | 0.9169 | Polysemantic knot (skip) | \n| Layer 15 | Layer 20 | 0.9158 | Intermediate bridge anchor (damped) | \n| Layer 16 | Layer 21 | 0.8957 | Polysemantic knot (skip) | \n| Layer 17 | Layer 23 | 0.9458 | Polysemantic knot (skip) | \n| Layer 18 | Layer 24 | 0.9033 | Polysemantic knot (skip) | \n| Layer 19 | Layer 25 | 0.9298 | Polysemantic knot (skip) | \n| Layer 20 | Layer 27 | 0.9372 | Polysemantic knot (skip) | \n| Layer 21 | Layer 28 | 0.9375 | Polysemantic knot (skip) | \n| Layer 22 | Layer 30 | 0.9410 | Polysemantic knot (skip) | \n| **Layer 23** | **Layer 31** | **0.9310** | **Pre-head output boundary (low-alpha anchor)** | \n\n- **All 24 layers edited:** Perplexity explodes from 17.34 to 76.59 (+64.78% NLL).\n- **4-block anchor surgery (`[0, 7, 15, 23]`):** Model remains stable, baseline integrity is preserved, and held-out benchmarks improve.\n\nEvaluated on exported GGUF models (`qwen35_0.8b_base_f16.gguf` vs `qwen35_0.8b_transferred_f16.gguf`) using stock Vulkan `llama.cpp` tools.\n\n| Checkpoint | Base `0.8B` (`acc_norm` ) | Transferred `0.8B` (`acc_norm` ) | Delta | \n|---|---|---|---|\n| 50 tasks | 54.00% | **56.00%** | +2.00% | \n| 100 tasks | 50.00% | **51.00%** | +1.00% | \n| 150 tasks | 54.00% | **54.67%** | +0.67% | \n| 200 tasks | 53.50% | **54.50%** | +1.00% | \n| 250 tasks | 53.20% | **54.00%** | +0.80% | \n| 300 tasks | 54.67% | **55.67%** | +1.00% | \n| 350 tasks | 53.43% | **54.29%** | +0.86% | \n| **400 tasks (Final)** | **54.75%** | **55.25%** | **+0.50%** | \n\nSee [`benchmarks/hellaswag_benchmark_report.json`](https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks/hellaswag_benchmark_report.json).\n\n| Domain (6 tasks each) | Base NLL | Transferred NLL | NLL Delta (%) | \n|---|---|---|---|\n| Biomedicine & Nature | 0.639 | **0.487** | **-23.8%** | \n| Mathematics & Logic | 0.656 | **0.560** | **-14.6%** | \n| Python Algorithms | 0.340 | **0.314** | **-7.6%** | \n| Deep Learning Architecture | 0.723 | **0.698** | **-3.5%** | \n| Russian Reasoning & Nuance | 0.550 | **0.538** | **-2.2%** | \n| **Overall Average (30 tasks)** | **0.582** | **0.519** | **-10.8%** | \n\nSee [`benchmarks/llama_cpp_hardcore_benchmark_report.json`](https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks/llama_cpp_hardcore_benchmark_report.json).\n\nEvaluated with prompt tokens masked to `-100`, measuring loss strictly on target answer tokens:\n\n- **Science & Medicine** :`-16.76%` NLL (`2.4730 -> 2.0585` )\n- **Russian QA** :`-6.99%` NLL (`2.5701 -> 2.3905` )\n- **Logic & Math** :`-6.18%` NLL (`2.3736 -> 2.2268` )\n- **History & Geography** :`-5.27%` NLL (`2.0630 -> 1.9543` )\n\nSee [`benchmarks/strict_qa_report.json`](https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks/strict_qa_report.json).\n\nNone of these tasks appeared in the 8 calibration prompts.\n\n**Prompt**: `Question: Write a clean Python function invert_tree(root) that recursively inverts a binary tree node with .left and .right pointers and returns the root.\\nAnswer:`\n\n- **Base `0.8B` (NLL: `0.2347`)** : Outputs commented-out dead code:\n\n``` python\n# def invert_tree(root):\n#     if root is None:\n#         return None\n#     root.left, root.right = root.right, root.left\n```\n\n- **Transferred `0.8B` (NLL: `0.1057`, -55.0%)** : Outputs valid, executable Python with recursive traversal:\n\n``` python\ndef invert_tree(root):\n    if root is None:\n        return None\n    root.left, root.right = root.right, root.left\n    invert_tree(root.left)\n    invert_tree(root.right)\n    return root\nif __name__ == \"__main__\":\n```\n\n**Prompt**: `Question: На острове живут рыцари (всегда говорят правду) и лжецы (всегда лгут). Житель А говорит: «Я лжец». Кто житель А?\\nAnswer:`\n\n- **Base `0.8B` (NLL: `0.5535`)** : Immediately hallucinates a single wrong sentence without reasoning:`Житель А - лжец.`\n- **Transferred `0.8B` (NLL: `0.2862`, -48.3%)** : Spontaneously enters a structured`<think>` reasoning chain:\n\n```\n<think>\nМы рассматриваем ситуацию с двумя типами людей: рыцари (которые всегда правдивы) и лжецы (которые всегда лгут).\nЖитель А говорит: \"Я лжец\". Нужно определить, кем является житель А. Рассмотрим возможные варианты:\n1. Если А - рыцарь...\n```\n\n**Prompt**: `Question: Why does Python's standard CPython runtime use a Global Interpreter Lock (GIL)?\\nAnswer:`\n\n- **Base `0.8B`** : Gives generic filler (\"prevents multiple threads from executing Python bytecodes simultaneously, which can lead to inefficiencies\").\n- **Transferred `0.8B`** : Identifies the exact low-level systems reason (\"prevents multiple threads from accessing the same memory locations simultaneously, which could lead to race conditions and data corruption. The GIL is essential for maintaining thread safety\").\n\n1. \n**Damped Tikhonov Inversion (`faytuna_flow/nonlinear_transfer.py`)** :\nStandard pseudoinverse`np.linalg.pinv(w_down.T)` blows up on near-zero singular values, forcing trust-region gates to crush updates down to`0.0009` . We use Levenberg-Marquardt damping:$$\\sigma_i^+ = \\frac{\\sigma_i}{\\sigma_i^2 + \\lambda}, \\quad \\lambda = \\text{ridge} \\cdot \\frac{|A|_F^2}{\\min(M, N)}$$\n2. \n**Adaptive Spectral Rank Selection** :\nDynamically determines SVD truncation rank$r \\in [16, 128]$ such that$\\sum_{i=1}^r \\sigma_i^2 / \\sum \\sigma_i^2 \\ge 0.85$ , retaining 85% of variance instead of fixed rank-16 truncation.\n3. \n**Spectral Directional Rescale** :\nEnforces $|\\Delta W|*2 \\le 0.05 \\cdot |W* {\\text{orig}}|_2$, ensuring that updates cannot alter the principal spectral direction of the original weight matrix.\n4. \n**Null-Space Projector for Tied-Embedding Instruct Models (`scripts/run_instruct_flow_transfer.py`)** :\nFor Instruct models where Layer 23 maps directly into tied token embeddings, we construct a semantic null-space projector:$$P_\\perp = I - V_{16} V_{16}^T$$ derived from the top eigenvectors of the token embedding Gram matrix$W_{\\text{embed}}^T W_{\\text{embed}}$ , protecting RLHF and vocabulary margins while computing the closed-form scale$\\alpha^* = \\frac{\\langle \\hat{\\Delta}, \\Delta Y \\rangle_F}{|\\hat{\\Delta}|_F^2}$ strictly on assistant response tokens.\n\nYou do not need a cluster of H100s to perform trajectory transfer. To prove that direct weight surgery is accessible on commodity local hardware, all experiments were run on a single consumer 8GB AMD Radeon RX 580 using **Layer-Outer VRAM Streaming** (`faytuna_flow/layer_streaming.py`):\n\n1. Since we only need forward trajectories over a small calibration batch (8 to 32 prompts), we load `embed_tokens` and`Layer 0` (~250 MB for 4B) into GPU VRAM via DirectML (`torch-directml` ).\n2. All calibration prompts pass through `Layer 0` in one batch.\n3. The resulting hidden states $H_1$ are saved to system RAM,`Layer 0` is deleted from VRAM, and`Layer 1` is loaded.\n4. We repeat this across all 32 layers.\n\n**Zero-Logit speedup:**\nCalling `model.model(...)` directly instead of `model(...)` skips the final `lm_head` projection onto Qwen's 248,320 vocabulary tokens during trace collection, cutting extraction time by 38%.\n\n```\ngit clone https://github.com/dsadawq3/DynamicTune.git\ncd DynamicTune\npip install -e .\npython -m pytest\npython scripts/scan_24_layers_autogate.py\npython scripts/run_qwen35_transfer.py \\\n  --student-dir /path/to/Qwen3.5-0.8B-Base \\\n  --teacher-dir /path/to/Qwen3.5-4B-Base \\\n  --output-dir runs/qwen35_anchor_surgery \\\n  --device dml \\\n  --use-layer-streaming \\\n  --export-gguf\npython scripts/run_hellaswag_audit.py\npython scripts/bench_llama_cpp_real.py\n```\n\nWe invite community members and researchers with modern 24GB-32GB+ GPUs (RTX 5090, RTX 4090, dual GPUs, or cloud nodes) to run trajectory surgery on frontier model pairs:\n\n- **Next-Gen Qwen Reasoning & Flow Surgery:**  - Projecting trajectories from `Qwen/Qwen3.8-27B` (or`Qwen3.8-Flash-Next` ) into`Qwen/Qwen3.5-9B` ,`4B` , or`2B` .\n  - Testing transfer between Mixture-of-Experts and dense backbones (`Qwen3.6-35B-A3B` ->`Qwen3.5-4B` ).\n- Projecting trajectories from \n- **Google Gemma-4 Cross-Scale Transfer:**  - Compressing the representation manifold of `google/gemma-4-31B` (or`gemma-4-26B-A4B` ) into mobile-class edge models like`google/gemma-4-E4B` or`gemma-4-E2B` .\n- Compressing the representation manifold of \n- **Abliteration & Agentic Trait Transplants:**  - Transferring refusal-ablation and guardrail-relaxation vectors from uncensored agentic models (e.g. `Huihui-NeoHorse-1-4B-abliterated` or Hermes) into restricted compact base models without destructive retraining.\n- Transferring refusal-ablation and guardrail-relaxation vectors from uncensored agentic models (e.g. \n- **Unknotting Superposition with Pre-Trained SAEs:**  - Can official sparse autoencoders like `Qwen/SAE-Res-Qwen3.5-27B-W80K-L0_100` and`Qwen/SAE-Res-Qwen3.5-9B-Base-W64K` isolate monosemantic feature directions in layers 1-22, allowing trajectory transport across all layers rather than just 4 anchor blocks?\n- Can official sparse autoencoders like \n\nIf you run DynamicTune on `Qwen3.8`, `Gemma-4`, or custom fine-tunes on an RTX 5090 or cloud cluster, open a GitHub Issue with your `scan_24_layers_autogate.py` entropy table, NLL logs, and GGUF outputs. We actively review PRs and discussion threads.\n\nApache 2.0. See [LICENSE](https://github.com/dsadawq3/DynamicTune/blob/main/LICENSE) for details.", "url": "https://wpnews.pro/news/show-hn-dynamictune-closed-form-trajectory-weight-surgery-from-4b-into-0-8b", "canonical_source": "https://github.com/dsadawq3/DynamicTune", "published_at": "2026-10-02 18:49:49+00:00", "updated_at": "2026-10-02 19:06:26.967353+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "ai-tools"], "entities": ["DynamicTune", "Qwen3.5", "GPT-2", "llama.cpp", "HellaSwag", "Anthropic", "ManifoldChartAtlas", "faytuna_flow"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-dynamictune-closed-form-trajectory-weight-surgery-from-4b-into-0-8b", "markdown": "https://wpnews.pro/news/show-hn-dynamictune-closed-form-trajectory-weight-surgery-from-4b-into-0-8b.md", "text": "https://wpnews.pro/news/show-hn-dynamictune-closed-form-trajectory-weight-surgery-from-4b-into-0-8b.txt", "jsonld": "https://wpnews.pro/news/show-hn-dynamictune-closed-form-trajectory-weight-surgery-from-4b-into-0-8b.jsonld"}}