Retrofitting language models to operate over bytes A Nature paper introduces 'byteification', a two-stage conversion procedure that retrofits existing subword-tokenized large language models into byte-level models with minimal extra training. The resulting byte-level models outperform earlier byte-level approaches and excel on character-level reasoning tasks while achieving practical inference speeds and reusing the existing ecosystem around the source model, removing a long-standing performance barrier to end-to-end byte-level language modelling. The authors argue subword tokenization obscures fine-grained information needed for scientific data such as computer code and biological sequences, and biases models toward English-centric vocabularies. Abstract Recent advances in artificial intelligence AI have largely been driven by large language models, deep neural networks that operate over discrete units called tokens. To represent text, most large language models use words or word fragments as the tokens, known as subword tokenization