cd /news/machine-learning/show-hn-dynamictune-closed-form-traj… · home › topics › machine-learning › article
[ARTICLE · art-144080] src=github.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B

DynamicTune, a Show HN project, reports that closed-form weight surgery can transfer trajectory dynamics from a 4B-parameter Qwen3.5 (d=2560) into a 0.8B student (d=1024) and from GPT-2 XL (d=1600) into GPT-2 small (d=768) without backpropagation or training runs. Editing all 24 student layers raised negative log-likelihood by 64.78%, but restricting the surgery to 4 anchor blocks (layers 0, 7, 15, 23) cut multi-domain held-out NLL by 10.8% and improved HellaSwag by 0.50% across 400 tasks on native Vulkan llama.cpp. The project attributes the failure of full-stack edits to high spectral entropy (>0.90) in intermediate layers 1-22 and provides reproducible benchmark logs in its benchmarks directory.

read11 min views1 publishedOct 2, 2026
Show HN: DynamicTune – Closed-form trajectory weight surgery from 4B into 0.8B
Image: Michielbdejong (auto-discovered)

Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths.

Tested across radically different model families: modern hybrid Qwen3.5 (4B with d=2560 -> 0.8B with d=1024) and notoriously fragile GPT-2 (XL with d=1600 -> small with d=768).

The prevailing consensus in deep learning is that transferring capabilities from a larger teacher model to a smaller student demands billions or trillions of tokens, massive synthetic dataset pipelines, and weeks of GPU cluster compute running token-level cross-entropy or KL divergence minimization.

DynamicTune challenges this dogma:

  1. A transformer stack is fundamentally a discrete dynamical system over depth: $h_{l+1} = h_l + f_l(h_l)$ .
  2. A capable teacher traces an informational velocity field through representation space.
  3. By aligning these trajectories through a local orthogonal Procrustes atlas and solving for closed-form weight updates in the student's MLP blocks, we can physically transfer teacher trajectory dynamics into the student without backpropagation or training runs.

Cross-architecture stability: We tested this on both modernQwen3.5 (SwiGLU, hybrid linear/sliding attention) and the notoriously fragileGPT-2 small (where Conv1D layers famously collapse or degrade into gibberish at the slightest weight disturbance). In both architectures, baseline language modeling integrity is preserved with well-behaved, bounded degradation margins, while target domain accuracy improves. 5. The spectral entropy discovery: Editing all 24 student layers destroys the model (+64.78% NLL) because intermediate layers (1-22) are high-entropy polysemantic knots (>0.90 spectral entropy). Restricting the surgery to4 anchor blocks (layers 0, 7, 15, and 23) avoids destructive interference, cuts multi-domain held-out NLL by**-10.8%** , and boostsHellaSwag (+0.50% across 400 tasks) on native Vulkanllama.cpp .

Raw reproducible benchmark logs: benchmarks/.

DynamicTune bridges established machine learning theory and mechanistic interpretability into an empirical weight surgery engine:

Residual connections allow layers to be viewed as Euler discretization steps of an underlying continuous ordinary differential equation

In DynamicTune, we do not view weights as static feature matrices. We treat the step

Concepts in large language models are represented as linear directions in representation space, and different models often learn linearly or orthogonally equivalent geometries up to rotation and scaling.

Because teacher and student models have different hidden dimensions (e.g. 2560 vs 1024), a single global orthogonal matrix cannot capture non-linear curvature across different semantic clusters. DynamicTune builds a piecewise local Procrustes atlas (ManifoldChartAtlas in faytuna_flow/manifold_charts.py): we cluster hidden states with K-Means into

Why did past attempts at layer-wise weight transfer fail? Anthropic's research into mechanistic interpretability showed that neural networks pack more features than they have dimensions via superposition, creating polysemantic neurons that activate on multiple unrelated concepts.

When a student model has only 1024 dimensions, intermediate layers (layers 1 to 22) are forced to operate in dense superposition. In DynamicTune, we compute the singular value distribution of the flow residual and calculate its normalized Shannon spectral entropy faytuna_flow/knots.py).

Layer 0 exhibits low entropy ($H = 0.7138$ ): clean, coherent semantic grounding. #

Layers 1-22 exhibit high entropy ($H \in [0.8957, 0.9669]$ ): chaotic superposition knots. Forcing a linear weight update here creates catastrophic interference and ruins the model. #

Layer 23 ($H = 0.9310$ ): pre-unembed boundary where features unpack toward vocabulary logits.

By discovering this entropy barrier, we learned that weight surgery must respect superposition boundaries: edit sparse anchor points, bypass the knots.

Instead of gradient descent, direct weight updates can be computed as closed-form linear projections that satisfy key-value associations.

DynamicTune extends this concept from individual fact-editing to depth-wise dynamical flow transport: we pull the projected trajectory deltas back through the SwiGLU MLP blocks via a damped Tikhonov pseudoinverse and rank-constrained SVD projections with explicit spectral trust-region bounds.

Running scripts/scan_24_layers_autogate.py across all layers of Qwen3.5-0.8B mapped to Qwen3.5-4B reveals why full-model transfer fails:

Student Layer Mapped Teacher Layer Multi-Chart Atlas Spectral Entropy Diagnosis
Layer 0 Layer 0 0.7138 Coherent semantic anchor (safe for surgery)
Layer 1 Layer 1 0.9669 Polysemantic knot (skip)
Layer 2 Layer 3 0.9419 Polysemantic knot (skip)
Layer 3 Layer 4 0.9032 Polysemantic knot (skip)
Layer 4 Layer 5 0.9572 Polysemantic knot (skip)
Layer 5 Layer 7 0.9250 Polysemantic knot (skip)
Layer 6 Layer 8 0.9388 Polysemantic knot (skip)
Layer 7 Layer 9 0.9384 Intermediate bridge anchor (damped)
Layer 8 Layer 11 0.9387 Polysemantic knot (skip)
Layer 9 Layer 12 0.9531 Polysemantic knot (skip)
Layer 10 Layer 13 0.9182 Polysemantic knot (skip)
Layer 11 Layer 15 0.9482 Polysemantic knot (skip)
Layer 12 Layer 16 0.9504 Polysemantic knot (skip)
Layer 13 Layer 17 0.9183 Polysemantic knot (skip)
Layer 14 Layer 19 0.9169 Polysemantic knot (skip)
Layer 15 Layer 20 0.9158 Intermediate bridge anchor (damped)
Layer 16 Layer 21 0.8957 Polysemantic knot (skip)
Layer 17 Layer 23 0.9458 Polysemantic knot (skip)
Layer 18 Layer 24 0.9033 Polysemantic knot (skip)
Layer 19 Layer 25 0.9298 Polysemantic knot (skip)
Layer 20 Layer 27 0.9372 Polysemantic knot (skip)
Layer 21 Layer 28 0.9375 Polysemantic knot (skip)
Layer 22 Layer 30 0.9410 Polysemantic knot (skip)
Layer 23 Layer 31 0.9310 Pre-head output boundary (low-alpha anchor)
  • All 24 layers edited: Perplexity explodes from 17.34 to 76.59 (+64.78% NLL).
  • 4-block anchor surgery ([0, 7, 15, 23]): Model remains stable, baseline integrity is preserved, and held-out benchmarks improve.

Evaluated on exported GGUF models (qwen35_0.8b_base_f16.gguf vs qwen35_0.8b_transferred_f16.gguf) using stock Vulkan llama.cpp tools.

Checkpoint Base 0.8B (acc_norm ) Transferred 0.8B (acc_norm ) Delta
50 tasks 54.00% 56.00% +2.00%
100 tasks 50.00% 51.00% +1.00%
150 tasks 54.00% 54.67% +0.67%
200 tasks 53.50% 54.50% +1.00%
250 tasks 53.20% 54.00% +0.80%
300 tasks 54.67% 55.67% +1.00%
350 tasks 53.43% 54.29% +0.86%
400 tasks (Final) 54.75% 55.25% +0.50%

See benchmarks/hellaswag_benchmark_report.json.

Domain (6 tasks each) Base NLL Transferred NLL NLL Delta (%)
Biomedicine & Nature 0.639 0.487 -23.8%
Mathematics & Logic 0.656 0.560 -14.6%
Python Algorithms 0.340 0.314 -7.6%
Deep Learning Architecture 0.723 0.698 -3.5%
Russian Reasoning & Nuance 0.550 0.538 -2.2%
Overall Average (30 tasks) 0.582 0.519 -10.8%

See benchmarks/llama_cpp_hardcore_benchmark_report.json.

Evaluated with prompt tokens masked to -100, measuring loss strictly on target answer tokens:

  • Science & Medicine :-16.76% NLL (2.4730 -> 2.0585 )
  • Russian QA :-6.99% NLL (2.5701 -> 2.3905 )
  • Logic & Math :-6.18% NLL (2.3736 -> 2.2268 )
  • History & Geography :-5.27% NLL (2.0630 -> 1.9543 )

See benchmarks/strict_qa_report.json.

None of these tasks appeared in the 8 calibration prompts.

Prompt: Question: Write a clean Python function invert_tree(root) that recursively inverts a binary tree node with .left and .right pointers and returns the root.\nAnswer:

  • Base 0.8B (NLL: 0.2347) : Outputs commented-out dead code:
  • Transferred 0.8B (NLL: 0.1057, -55.0%) : Outputs valid, executable Python with recursive traversal:
def invert_tree(root):
    if root is None:
        return None
    root.left, root.right = root.right, root.left
    invert_tree(root.left)
    invert_tree(root.right)
    return root
if __name__ == "__main__":

Prompt: Question: На острове живут рыцари (всегда говорят правду) и лжецы (всегда лгут). Житель А говорит: «Я лжец». Кто житель А?\nAnswer:

  • Base 0.8B (NLL: 0.5535) : Immediately hallucinates a single wrong sentence without reasoning:Житель А - лжец.
  • Transferred 0.8B (NLL: 0.2862, -48.3%) : Spontaneously enters a structured<think> reasoning chain:
<think>
Мы рассматриваем ситуацию с двумя типами людей: рыцари (которые всегда правдивы) и лжецы (которые всегда лгут).
Житель А говорит: "Я лжец". Нужно определить, кем является житель А. Рассмотрим возможные варианты:
1. Если А - рыцарь...

Prompt: Question: Why does Python's standard CPython runtime use a Global Interpreter Lock (GIL)?\nAnswer:

  • Base 0.8B : Gives generic filler ("prevents multiple threads from executing Python bytecodes simultaneously, which can lead to inefficiencies").
  • Transferred 0.8B : Identifies the exact low-level systems reason ("prevents multiple threads from accessing the same memory locations simultaneously, which could lead to race conditions and data corruption. The GIL is essential for maintaining thread safety").

Damped Tikhonov Inversion (faytuna_flow/nonlinear_transfer.py) : Standard pseudoinversenp.linalg.pinv(w_down.T) blows up on near-zero singular values, forcing trust-region gates to crush updates down to0.0009 . We use Levenberg-Marquardt damping:$$\sigma_i^+ = \frac{\sigma_i}{\sigma_i^2 + \lambda}, \quad \lambda = \text{ridge} \cdot \frac{|A|F^2}{\min(M, N)}$$ 2. Adaptive Spectral Rank Selection : Dynamically determines SVD truncation rank$r \in [16, 128]$ such that$\sum{i=1}^r \sigma_i^2 / \sum \sigma_i^2 \ge 0.85$ , retaining 85% of variance instead of fixed rank-16 truncation. 3. Spectral Directional Rescale : Enforces $|\Delta W|2 \le 0.05 \cdot |W {\text{orig}}|2$, ensuring that updates cannot alter the principal spectral direction of the original weight matrix. 4. Null-Space Projector for Tied-Embedding Instruct Models (scripts/run_instruct_flow_transfer.py) : For Instruct models where Layer 23 maps directly into tied token embeddings, we construct a semantic null-space projector:$$P\perp = I - V_{16} V_{16}^T$$ derived from the top eigenvectors of the token embedding Gram matrix$W_{\text{embed}}^T W_{\text{embed}}$ , protecting RLHF and vocabulary margins while computing the closed-form scale$\alpha^* = \frac{\langle \hat{\Delta}, \Delta Y \rangle_F}{|\hat{\Delta}|_F^2}$ strictly on assistant response tokens.

You do not need a cluster of H100s to perform trajectory transfer. To prove that direct weight surgery is accessible on commodity local hardware, all experiments were run on a single consumer 8GB AMD Radeon RX 580 using Layer-Outer VRAM Streaming (faytuna_flow/layer_streaming.py):

  1. Since we only need forward trajectories over a small calibration batch (8 to 32 prompts), we load embed_tokens andLayer 0 (~250 MB for 4B) into GPU VRAM via DirectML (torch-directml ).
  2. All calibration prompts pass through Layer 0 in one batch.
  3. The resulting hidden states $H_1$ are saved to system RAM,Layer 0 is deleted from VRAM, andLayer 1 is loaded.
  4. We repeat this across all 32 layers.

Zero-Logit speedup: Calling model.model(...) directly instead of model(...) skips the final lm_head projection onto Qwen's 248,320 vocabulary tokens during trace collection, cutting extraction time by 38%.

git clone https://github.com/dsadawq3/DynamicTune.git
cd DynamicTune
pip install -e .
python -m pytest
python scripts/scan_24_layers_autogate.py
python scripts/run_qwen35_transfer.py \
  --student-dir /path/to/Qwen3.5-0.8B-Base \
  --teacher-dir /path/to/Qwen3.5-4B-Base \
  --output-dir runs/qwen35_anchor_surgery \
  --device dml \
  --use-layer-streaming \
  --export-gguf
python scripts/run_hellaswag_audit.py
python scripts/bench_llama_cpp_real.py

We invite community members and researchers with modern 24GB-32GB+ GPUs (RTX 5090, RTX 4090, dual GPUs, or cloud nodes) to run trajectory surgery on frontier model pairs:

  • Next-Gen Qwen Reasoning & Flow Surgery: - Projecting trajectories from Qwen/Qwen3.8-27B (orQwen3.8-Flash-Next ) intoQwen/Qwen3.5-9B ,4B , or2B .
    • Testing transfer between Mixture-of-Experts and dense backbones (Qwen3.6-35B-A3B ->Qwen3.5-4B ).
  • Projecting trajectories from
  • Google Gemma-4 Cross-Scale Transfer: - Compressing the representation manifold of google/gemma-4-31B (orgemma-4-26B-A4B ) into mobile-class edge models likegoogle/gemma-4-E4B orgemma-4-E2B .
  • Compressing the representation manifold of
  • Abliteration & Agentic Trait Transplants: - Transferring refusal-ablation and guardrail-relaxation vectors from uncensored agentic models (e.g. Huihui-NeoHorse-1-4B-abliterated or Hermes) into restricted compact base models without destructive retraining.
  • Transferring refusal-ablation and guardrail-relaxation vectors from uncensored agentic models (e.g.
  • Unknotting Superposition with Pre-Trained SAEs: - Can official sparse autoencoders like Qwen/SAE-Res-Qwen3.5-27B-W80K-L0_100 andQwen/SAE-Res-Qwen3.5-9B-Base-W64K isolate monosemantic feature directions in layers 1-22, allowing trajectory transport across all layers rather than just 4 anchor blocks?
  • Can official sparse autoencoders like

If you run DynamicTune on Qwen3.8, Gemma-4, or custom fine-tunes on an RTX 5090 or cloud cluster, open a GitHub Issue with your scan_24_layers_autogate.py entropy table, NLL logs, and GGUF outputs. We actively review PRs and discussion threads.

Apache 2.0. See LICENSE for details.

── more in #machine-learning 4 stories · sorted by recency
── more on @dynamictune 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-dynamictune-…] indexed:0 read:11min 2026-10-02 · —