Critic-free Reward-Integrated Self-distillation Policy Optimization
-
Updated
Aug 2, 2026 - Python
Critic-free Reward-Integrated Self-distillation Policy Optimization
LLM Post-Training, RLHF, PPO, DPO, etc
Coordinated AI systems evolving together.
Add a description, image, and links to the post-training-learning topic page so that developers can more easily learn about it.
To associate your repository with the post-training-learning topic, visit your repo's landing page and select "manage topics."