How Anthropic trained Fable 5 => by analysing its reasoning traces
Anthropic trained its Fable-5 model by analyzing its reasoning traces, using a post-training process of reinforcement learning, synthetic data generation, and self-distillation. The model excelled at cybersecurity tasks …