Can Language Models Remember What They Learn?
Post-training methods (RLVR, On-policy distillation) are Episode-local Language models are getting better at learning from feedback during post-training. In reinforcement learning with verifiable rewards (RLVR), a model tries a problem, a verifier checks the answer, and the policy is updated based on the scalar reward. Recent self-distillation methods go further by using feedback or successful […]
Can Language Models Remember What They Learn? Read More »







