Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale A study at 1.7B-class scale found that a shared Transformer can learn language modeling from fixed token identities rather than a trainable input embedding table, according to the paper's headline claim. The research compares fixed minimal token codes against token-specific adjustable vectors to test whether per-token parameterization is required for substantial capability. A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare t