On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
-
Updated
Aug 31, 2026 - Python
On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
Official code for "Self-Distilled Agentic Reinforcement Learning"
Source code of paper "RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation"
TTPO: Test-Time Policy Optimization
A PyTorch-based, end-to-end energy time-series forecasting demo with LSTM/GRU models, strong seasonal baselines, residual learning, and a CLI for reproducible train/eval runs on OPSD (with run logging + notebooks).
To associate your repository with the opsd topic, visit your repo's landing page and select "manage topics."