Learning to Learn a Language Researchers introduced the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior, which learns to predict real text in context with frozen weights despite never having seen a word of any real language. The model adapts to a real-text prefix without any weight updates, according to the work. We present the Prior-Fitted Language Model PFLM , a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. E