Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B DynamicTune, a Show HN project, reports that closed-form weight surgery can transfer trajectory dynamics from a 4B-parameter Qwen3.5 (d=2560) into a 0.8B student (d=1024) and from GPT-2 XL (d=1600) into GPT-2 small (d=768) without backpropagation or training runs. Editing all 24 student layers raised negative log-likelihood by 64.78%, but restricting the surgery to 4 anchor blocks (layers 0, 7, 15, 23) cut multi-domain held-out NLL by 10.8% and improved HellaSwag by 0.50% across 400 tasks on native Vulkan llama.cpp. The project attributes the failure of full-stack edits to high spectral entropy (>0.90) in intermediate layers 1-22 and provides reproducible benchmark logs in its benchmarks directory. Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths. Tested across radically different model families: modern hybrid Qwen3.5 4B with d=2560 - 0.8B with d=1024 and notoriously fragile GPT-2 XL with d=1600 - small with d=768 . The prevailing consensus in deep learning is that transferring capabilities from a larger teacher model to a smaller student demands billions or trillions of tokens, massive synthetic dataset pipelines, and weeks of GPU cluster compute running token-level cross-entropy or KL divergence minimization. DynamicTune challenges this dogma: 1. A transformer stack is fundamentally a discrete dynamical system over depth: $h {l+1} = h l + f l h l $ . 2. A capable teacher traces an informational velocity field through representation space. 3. By aligning these trajectories through a local orthogonal Procrustes atlas and solving for closed-form weight updates in the student's MLP blocks, we can physically transfer teacher trajectory dynamics into the student without backpropagation or training runs. 4. Cross-architecture stability: We tested this on both modern Qwen3.5 SwiGLU, hybrid linear/sliding attention and the notoriously fragile GPT-2 small where Conv1D layers famously collapse or degrade into gibberish at the slightest weight disturbance . In both architectures, baseline language modeling integrity is preserved with well-behaved, bounded degradation margins, while target domain accuracy improves. 5. The spectral entropy discovery: Editing all 24 student layers destroys the model +64.78% NLL because intermediate layers 1-22 are high-entropy polysemantic knots 0.90 spectral entropy . Restricting the surgery to 4 anchor blocks layers 0, 7, 15, and 23 avoids destructive interference, cuts multi-domain held-out NLL by -10.8% , and boosts HellaSwag +0.50% across 400 tasks on native Vulkan llama.cpp . Raw reproducible benchmark logs: benchmarks/ https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks . DynamicTune bridges established machine learning theory and mechanistic interpretability into an empirical weight surgery engine: Residual connections allow layers to be viewed as Euler discretization steps of an underlying continuous ordinary differential equation - Chen et al., 2018: Neural Ordinary Differential Equations arXiv:1806.07366 https://arxiv.org/abs/1806.07366 - Lu et al., 2017: Beyond Finite Layer Neural Networks: Bridging Deep Architectures and Numerical Differential Equations arXiv:1710.10121 https://arxiv.org/abs/1710.10121 - Sander et al., 2022: Residual Neural Networks as Approximations of Ordinary Differential Equations arXiv:2202.10512 https://arxiv.org/abs/2202.10512 In DynamicTune, we do not view weights as static feature matrices. We treat the step Concepts in large language models are represented as linear directions in representation space, and different models often learn linearly or orthogonally equivalent geometries up to rotation and scaling. - Park et al., 2023: The Linear Representation Hypothesis and the Geometry of Large Language Models arXiv:2311.03658 https://arxiv.org/abs/2311.03658 - Kornblith et al., 2019: Similarity of Neural Network Representations Revisited arXiv:1905.00414 https://arxiv.org/abs/1905.00414 - Ding et al., 2021: Grounding Representation Similarity with Statistical Mechanics arXiv:2106.11561 https://arxiv.org/abs/2106.11561 Because teacher and student models have different hidden dimensions e.g. 2560 vs 1024 , a single global orthogonal matrix cannot capture non-linear curvature across different semantic clusters. DynamicTune builds a piecewise local Procrustes atlas ManifoldChartAtlas in faytuna flow/manifold charts.py : we cluster hidden states with K-Means into Why did past attempts at layer-wise weight transfer fail? Anthropic's research into mechanistic interpretability showed that neural networks pack more features than they have dimensions via superposition, creating polysemantic neurons that activate on multiple unrelated concepts. - Elhage et al., 2022: Toy Models of Superposition arXiv:2209.10652 https://arxiv.org/abs/2209.10652 - Bricken et al., 2023: Towards Monosemanticity: Decomposing Language Models With Dictionary Learning https://transformer-circuits.pub/2023/monosemantic-features/index.html When a student model has only 1024 dimensions, intermediate layers layers 1 to 22 are forced to operate in dense superposition. In DynamicTune, we compute the singular value distribution of the flow residual and calculate its normalized Shannon spectral entropy faytuna flow/knots.py . - Layer 0 exhibits low entropy $H = 0.7138$ : clean, coherent semantic grounding. - Layers 1-22 exhibit high entropy $H \in 0.8957, 0.9669 $ : chaotic superposition knots. Forcing a linear weight update here creates catastrophic interference and ruins the model. - Layer 23 $H = 0.9310$ : pre-unembed boundary where features unpack toward vocabulary logits. By discovering this entropy barrier, we learned that weight surgery must respect superposition boundaries: edit sparse anchor points, bypass the knots. Instead of gradient descent, direct weight updates can be computed as closed-form linear projections that satisfy key-value associations. - Meng et al., 2022: Locating and Editing Factual Associations in GPT ROME, arXiv:2202.05262 https://arxiv.org/abs/2202.05262 - Meng et al., 2022: Mass-Editing Memory in a Transformer MEMIT, arXiv:2210.07229 https://arxiv.org/abs/2210.07229 DynamicTune extends this concept from individual fact-editing to depth-wise dynamical flow transport: we pull the projected trajectory deltas back through the SwiGLU MLP blocks via a damped Tikhonov pseudoinverse and rank-constrained SVD projections with explicit spectral trust-region bounds. Running scripts/scan 24 layers autogate.py across all layers of Qwen3.5-0.8B mapped to Qwen3.5-4B reveals why full-model transfer fails: | Student Layer | Mapped Teacher Layer | Multi-Chart Atlas Spectral Entropy | Diagnosis | |---|---|---|---| | Layer 0 | Layer 0 | 0.7138 | Coherent semantic anchor safe for surgery | | Layer 1 | Layer 1 | 0.9669 | Polysemantic knot skip | | Layer 2 | Layer 3 | 0.9419 | Polysemantic knot skip | | Layer 3 | Layer 4 | 0.9032 | Polysemantic knot skip | | Layer 4 | Layer 5 | 0.9572 | Polysemantic knot skip | | Layer 5 | Layer 7 | 0.9250 | Polysemantic knot skip | | Layer 6 | Layer 8 | 0.9388 | Polysemantic knot skip | | Layer 7 | Layer 9 | 0.9384 | Intermediate bridge anchor damped | | Layer 8 | Layer 11 | 0.9387 | Polysemantic knot skip | | Layer 9 | Layer 12 | 0.9531 | Polysemantic knot skip | | Layer 10 | Layer 13 | 0.9182 | Polysemantic knot skip | | Layer 11 | Layer 15 | 0.9482 | Polysemantic knot skip | | Layer 12 | Layer 16 | 0.9504 | Polysemantic knot skip | | Layer 13 | Layer 17 | 0.9183 | Polysemantic knot skip | | Layer 14 | Layer 19 | 0.9169 | Polysemantic knot skip | | Layer 15 | Layer 20 | 0.9158 | Intermediate bridge anchor damped | | Layer 16 | Layer 21 | 0.8957 | Polysemantic knot skip | | Layer 17 | Layer 23 | 0.9458 | Polysemantic knot skip | | Layer 18 | Layer 24 | 0.9033 | Polysemantic knot skip | | Layer 19 | Layer 25 | 0.9298 | Polysemantic knot skip | | Layer 20 | Layer 27 | 0.9372 | Polysemantic knot skip | | Layer 21 | Layer 28 | 0.9375 | Polysemantic knot skip | | Layer 22 | Layer 30 | 0.9410 | Polysemantic knot skip | | Layer 23 | Layer 31 | 0.9310 | Pre-head output boundary low-alpha anchor | - All 24 layers edited: Perplexity explodes from 17.34 to 76.59 +64.78% NLL . - 4-block anchor surgery 0, 7, 15, 23 : Model remains stable, baseline integrity is preserved, and held-out benchmarks improve. Evaluated on exported GGUF models qwen35 0.8b base f16.gguf vs qwen35 0.8b transferred f16.gguf using stock Vulkan llama.cpp tools. | Checkpoint | Base 0.8B acc norm | Transferred 0.8B acc norm | Delta | |---|---|---|---| | 50 tasks | 54.00% | 56.00% | +2.00% | | 100 tasks | 50.00% | 51.00% | +1.00% | | 150 tasks | 54.00% | 54.67% | +0.67% | | 200 tasks | 53.50% | 54.50% | +1.00% | | 250 tasks | 53.20% | 54.00% | +0.80% | | 300 tasks | 54.67% | 55.67% | +1.00% | | 350 tasks | 53.43% | 54.29% | +0.86% | | 400 tasks Final | 54.75% | 55.25% | +0.50% | See benchmarks/hellaswag benchmark report.json https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks/hellaswag benchmark report.json . | Domain 6 tasks each | Base NLL | Transferred NLL | NLL Delta % | |---|---|---|---| | Biomedicine & Nature | 0.639 | 0.487 | -23.8% | | Mathematics & Logic | 0.656 | 0.560 | -14.6% | | Python Algorithms | 0.340 | 0.314 | -7.6% | | Deep Learning Architecture | 0.723 | 0.698 | -3.5% | | Russian Reasoning & Nuance | 0.550 | 0.538 | -2.2% | | Overall Average 30 tasks | 0.582 | 0.519 | -10.8% | See benchmarks/llama cpp hardcore benchmark report.json https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks/llama cpp hardcore benchmark report.json . Evaluated with prompt tokens masked to -100 , measuring loss strictly on target answer tokens: - Science & Medicine : -16.76% NLL 2.4730 - 2.0585 - Russian QA : -6.99% NLL 2.5701 - 2.3905 - Logic & Math : -6.18% NLL 2.3736 - 2.2268 - History & Geography : -5.27% NLL 2.0630 - 1.9543 See benchmarks/strict qa report.json https://github.com/dsadawq3/DynamicTune/blob/main/benchmarks/strict qa report.json . None of these tasks appeared in the 8 calibration prompts. Prompt : Question: Write a clean Python function invert tree root that recursively inverts a binary tree node with .left and .right pointers and returns the root.\nAnswer: - Base 0.8B NLL: 0.2347 : Outputs commented-out dead code: python def invert tree root : if root is None: return None root.left, root.right = root.right, root.left - Transferred 0.8B NLL: 0.1057 , -55.0% : Outputs valid, executable Python with recursive traversal: python def invert tree root : if root is None: return None root.left, root.right = root.right, root.left invert tree root.left invert tree root.right return root if name == " main ": Prompt : Question: На острове живут рыцари всегда говорят правду и лжецы всегда лгут . Житель А говорит: «Я лжец». Кто житель А?\nAnswer: - Base 0.8B NLL: 0.5535 : Immediately hallucinates a single wrong sentence without reasoning: Житель А - лжец. - Transferred 0.8B NLL: 0.2862 , -48.3% : Spontaneously enters a structured