A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
Source code of paper "RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation"