cd /news/large-language-models/ninaxander-feasibility-and-limits-of… · home › topics › large-language-models › article
[ARTICLE · art-142987] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space

A study posted to arXiv (2609.38261v1) introduced NinaXander, a method that connects layers of frozen language models from different architecture families through a single trained shared-latent adapter, composing the recurrent RWKV-4-Raven-7B with the Transformer-based Tulu-Pythia-6.9b. The configuration combining the first 5 layers of Pythia with the remaining 27 layers of RWKV cut the Transformer key-value cache by 84.4% with accuracy not significantly different from RWKV alone, but no composed model matched the parent model Pythia on multiple-choice accuracy and language-modeling performance dropped sharply on WikiText. The authors report that intermediate-representation correspondence held only in one favorable case — shared tokenizer, same depth and same hidden width — and does not show the models share a general semantic space.

by read1 min views1 publishedOct 1, 2026

arXiv:2609.38261v1 Announce Type: new Abstract: In this paper we propose NinaXander, a series of composed language models obtained by connecting layers of frozen language models from different architecture families with a single trained shared-latent adapter. A composed model runs the first layers of one model, converts the resulting intermediate representation once with the adapter, and then runs the remaining layers of the other model. Once the adapter is trained, several composed models that connect at different layers are obtained without retraining. Using the recurrent RWKV-4-Raven-7B and the Transformer-based Tulu-Pythia-6.9b, abbreviated as RWKV and Pythia, this study examines whether frozen models from different families can be recombined post hoc. The composed models answered multiple-choice questions, and those whose generations we examined produced syntactically well-formed text. The configuration that combines the first 5 layers of Pythia with the remaining 27 layers of RWKV reduced the Transformer key-value (KV) cache by 84.4% with accuracy not significantly different from that of RWKV alone. In multiple-choice accuracy, however, no composed model matched the parent model Pythia, and language-modeling performance decreased sharply on WikiText, a corpus of Wikipedia articles outside the training domain. The correspondence between intermediate representations was also obtained in one favorable case, with a shared tokenizer, the same depth, and the same hidden width, and does not show that the models share a general semantic space.

── more in #large-language-models 4 stories · sorted by recency
── more on @ninaxander 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ninaxander-feasibili…] indexed:0 read:1min 2026-10-01 · —