14:31
2026-10-09
gilesthomas.com
large-language-models
Fun with low-rank vocab matrices (and a bonus test loss reduction?)
Training GPT-2 small-style models with low-rank factorised embeddings on the input side alone reduced test loss, while applying the trick to both embeddings and the output head raised loss by 0.07, 0.…