Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 16 days ago • 98