Probability of an LLM Trajectory: Starting from the Chain Rule
A technical explainer derives the probability of a large language model trajectory from first principles using the chain rule, factorizing p(τ|θ) into the initial state distribution p(s0), policy terms πθ(at|st), and env…