Why do OpenAI's GPT-2 weights beat mine? Part five: data quality
Independent LLM-from-scratch experiments found that OpenAI's GPT-2 small (124M parameters) consistently outperformed the author's 163M-parameter models on an instruction fine-tuning task adapted from …