
When LLMs Distill On-Policy
AwesomeOPD is an awesome list summarising open-source repositories and papers for training LLMs (and VLMs / agents / draft models) with On-Policy Distillation (OPD) and On-Policy Self-Distillation (OPSD):
- 🎯 OPD = C1 + C2.
C1: student samples its own trajectoriesy ~ π_student(·|x)during training.C2: teacher provides per-token / sequence supervision on those student samples. Methods that only partially satisfy are flagged in 📝 Strictness notes per section. - 🪞 OPSD = special case where teacher is the same model, conditioned on privileged context (verified trace / answer / "be concise" prefix / longer context) or an earlier checkpoint.
- 🚀 Each entry is annotated along four design axes — teacher source (external · same model with privileged context · earlier checkpoint · multi-teacher · discriminator), supervision signal (logits / top-k / sequence reward / verbal score / discriminator / verifier / feature), rollout consumption (all / selected / truncated / replaced / as PG samples), and pipeline slot (cold-start / mid / RL-replacement / inside-RL / inter-stage / compression / continual-anchor).
- ⚠️ Built by reading paper PDFs, project pages, and source code with LLM coding agents; manually reviewed but errors possible. PRs welcome.
- 📌 If you find this repository helpful for your research, please cite it via the "Cite this repository" button in the right sidebar of the GitHub page.
- 📅 Last updated: 2026-08-26
Taxonomy:
- 📚 Surveys, Foundations & Position Papers — meta-references and seed papers (GKD, MiniLLM, Thinking Machines blog, Tencent / THUNLP surveys)
- 🔬 White-Box — logit-based OPD on student rollouts with an external teacher
- 🎭 Black-Box — discriminator / verbal / preference, no teacher logits
- ♻️ OPSD — privileged-context self-distillation (same model, different conditioning)
- 🔁 Iterative Self-Bootstrapping — same model as previous-checkpoint teacher
- 🤝 OPD-RL Hybrids — inside-RL OPD: KL-as-reward, RL+OPD fusion
- 🧠 Reasoning / 🖼️ Multimodal / 🤖 Agent & Embodied — by application; cuts across all teacher-source categories
- ⚡ Speculative-Decoding Distillation — drafter distillation; "student" is a draft model
- 🛠️ Frameworks & Toolkits — what to actually run
- 🏭 Industrial / Production Reports — what the labs ship
Shorthand: FKL = forward KL · RKL = reverse KL · JSD = Jensen–Shannon · Skew-KL / AKL = skewed / adaptive KL · 📄 paper-only = no public code yet.
Updates
📢 click to expand
- 2026-08-26 — monthly sweep of 2026-07-20 → 2026-08-26: 107 new entries across every section, each read and then independently re-verified. Two passes: seven per-section sweeps proposed candidates against C1+C2, then ten adversarial verifiers re-checked every accepted entry against the arXiv abstract or HTML full text, attacking rather than confirming it. No entry survived on a proposer's word alone.
- Headline numbers: OPSD +33 (the largest single-month addition the list has seen), Multimodal +27, White-Box +21, Agent +14, OPD-RL +10, Industrial +6, Surveys +6, Black-Box / SpecDec / Frameworks +1 each. Reasoning and Iterative Self-Bootstrapping had no qualifying new work. Zero duplicates against the 275 arXiv IDs already listed.
- Rejected on the bar (5, after full-text reads): DreamMimic (2608.22278) — "imitation-dominant", teacher-driven rollout fraction, the FA-OPD failure pattern; SA-OPSD (2607.17136) — BC onto a verifier-corrected action, and dated one day outside the window; DiffusionOPSD (2608.24646) — frozen behaviour policy, reward-gradient targets, no privileged context. REOPD (2608.11698) and CROP (2608.13387) were both moved out of OPD-RL into White-Box: a PPO-shaped surrogate is not an RL loop when nothing but teacher log-probabilities enters the objective — the same reading under which G-OPD already sits in White-Box.
- Metadata corrections found by verification: 41 Org cells the sweep had marked "unstated" were in fact stated in the arXiv HTML full text (among them Apple for Rubrics as Privileged Information, NIO for LS-MOPD, ByteDance for OPLD; StreamOPD's "unix-ai-lab" turned out to be a GitHub Pages slug, not an institution). 7 titles were wrong — five appended a parenthetical acronym the paper never uses, and two (FP-OPD, SPIRAL) were outright mis-stated. 3 entries marked
📄 paper-onlyhave live repos (LOPD, SSPO, Veritas++), and Motif 3's weights turned out to be openly released. Only one "unstated" Org survived scrutiny (PCD, 2607.28336). - New acronym collisions (6): SMOPD ×2, RP-OPSD ×2, TA-OPD ×2, SPOT ×2, REOPD vs. ReOPD/REOPOLD, OPD-V vs. OPSD-V — all documented in the relevant sections' strictness notes.
- The month's through-line: all six new Surveys entries are diagnostic or negative, and they converge independently on the same conclusion — the privileged information is doing less work than assumed. OP²SD swaps the reference for an unrelated problem's solution and stays competitive; Privileged Likelihood Is Not Automatically Value finds hindsight scoring near-random at AUC 0.505 with an outcome-only baseline beating every token-score variant; Test-Time Scaling argues OPD narrows the sampling budget rather than raising the ceiling and names it "illusory distillation". Read alongside the existing Rethinking OPSD for Thinking Models (2607.05184).
- 2026-07-23 (c) — independent re-audit of all 191 entries added in (a) and (b). Every entry was re-read against its PDF by a second reviewer with no access to the first pass's reasoning; 189/191 verdicts were reproduced (98.9%).
- Removed as out of scope (2): CollectionLoRA (2605.25378) — single-step LoRA image editing, no sequence model; FA-OPD (2605.27095) — MLP policy on Gym/D4RL continuous control, no language or generative-sequence component. The flow-matching entries that are retained (DiffusionOPD, D-OPSD, OPSD-V, dOPSD, AnyFlow, WMSD, π-Flow, Qwen-Image-2.0-RL) generate a supervised trajectory; these two do not.
- Corrected: a transcription error in Multi-Rollout OPD (AIME25 mean@8 41.21 → 25.41); a misattributed ablation in OPSD Compresses RLVR (−26.8% → −29.22% at −1.03 pp); conflated settings in ADWIN; a worst-case-only figure in GeoSD; teacher count in MAD-OPD; a loose "doubles" in Tool-Call Boundary Drift.
- Metadata: an unverifiable venue tag on Revisiting OPD; four wrong or over-specific Org cells (HPD listed a GitHub username; Draft-OPD listed an affiliation absent from the paper; CoPD "JD Explore" → JD.COM; EffOPD "Tencent Hunyuan" → Tencent); ten Date cells normalised to the arXiv-ID month; ShortOPD refiled from White-Box to OPSD (its teacher is its own pre-compression checkpoint, not an external model); two new acronym collisions documented (PBSD ×2, COPD/CoPD); all badge counts recomputed.
- Checked and confirmed correct: the eight cross-listed arXiv IDs are the intended cross-section duplicates;
HJSang/OPSD_OnPolicyDistillationgenuinely hosts three separate papers (PACED, TIP, Sparse-to-Dense); no dead repo links.
- 2026-07-23 (b) — backfill sweep of 2026-04-25 → 2026-06-20: 152 papers read in full, 120 added, 32 rejected. This period predates the previous update and had never been covered, so the list was missing roughly half the field's output during its busiest months. Highlights: Decoupling KL and Trajectories (the cleanest formal taxonomy of SFT/DAgger/offline-RL/OPD), Apple's Unmasking OPD, three independent parameter-geometry studies, the prefix-truncation family (Prune-OPD, ESR, ADWIN, Truncated OPD), cross-tokenizer OPD (SimCT, Tokenizer Barrier), Meta's logit-free OmniOPD, and two production reports with genuine MOPD (Kwai Keye-VL-2.0, Nemotron 3 Ultra). New acronym-collision and contested-finding notes added to the relevant strictness sections. Backfill entries carry their mechanism summary inline in the table rather than a separate technical-details row.
- 2026-07-23 (a) — large sweep (~60 entries), every paper PDF read and every repo link checked.
- Surveys: Formula-Driven Survey, NAIL, When Does Online IL Help, Demystifying OPD
- White-Box: IW-OPD, DEAR, ReNIO, PG-OPD, SEAD, DOPD, MOPD (Xiaomi), Blockwise Drift Gating, Direct-OPD, OPD², TOP-D, COPD, ShortOPD, TOPD
- OPSD: CaOPD, Vision-OPD, D-OPSD, PW-OPSD, OPSD Predictive Law, SDSD Diversity, PHF, Visual-OPSD, Denser ≠ Better, Purified OPSD, DemoPSD, Rethinking OPSD, GeoSD, AD-OPSD, dOPSD, PromptSD; BIRD → Iterative Self-Bootstrapping
- OPD-RL: OPID, ATOD, PASS, CRAFT, UCOB, DRIFT, RG-OPD, CPO, Distilled RL, CriPO, H²SD, CADENCE
- Multimodal: DiffusionOPD, V-Zero, H-OPD, OPSD-V, Med-OPD, OPD-IAD
- Agent: KbSD, Two-Phase Distillation, SEED, ReOPD, UI-MOPD, TurnOPD, FA-SD Failure, Tool-Call Boundary Drift, ROAD-VLA
- SpecDec: TIGER, AdaFlash · Frameworks: Spider, AsyncOPD, EasyOPD
- Industrial: Solar Open 2, KAT-Coder-V2.5, OvisOCR2, Mach-Mind-4-Flash, Agents-A1, Qwen-Image-2.0-RL, Audex
- Maintenance: TCOD upgraded to its released repo (COLM 2026); Draft-OPD relinked to the canonical
Simplified-Reasoningorg; venue tags added to HPD (ICML 2026) and Revisiting OPD (COLM 2026); all badge counts recomputed (several were stale) - Verified and excluded: TREK, REGEN, A Few Teacher Steps, Prefix-GRPO, HyperDFlash, SOPD-SocialNav, DataFlex-RL — each fails C1 or C2 on a full read; rationale in the relevant section's strictness notes
- 2026-06-23 — add FiRe-OPD (White-Box; hard trajectory filtering + soft token reweighting)
- 2026-06-19 — add d-OPSD, RLCSD, SSOPD (OPSD); TGPO, OPD+ (OPD-RL); Flow-OPD, Decomposed-OPD/VGS (Multimodal); ROPD (Black-Box); BRTS, OPRD (White-Box); Draft-OPD (SpecDec); SDAR (Agent); The Many Faces of OPD (Surveys)
- 2026-06-13 — add TRD (White-Box; trajectory-level refinement) + Li Jiang's OPD reflection blog
- 2026-06-08 — add EMPO² (OPSD; cross-listed into Agent)
- 2026-06-07 — add SPOT
- 2026-05-18 — add COPSD, MSD
- 2026-05-15 — add TCOD, Healthcare AI GYM, HyperEyes (and cross-list Skill-SD into Agent)
- 2026-05-14 — add CORD
- 2026-05-13 — add Uni-OPD; add HY-MT, Baichuan-M3, KAT-Coder-V2, HY-Embodied, Qwen3.5-Omni
- 2026-05-11 — add Skill-SD
- 2026-04-30 — add π-Play
- 2026-04-29 — add SD-Zero, Why Does Self-Distillation (Sometimes) Degrade Reasoning?
- 2026-04-28 — initial release; add NPO, VLA-OPD, KDFlow, HPD, DeepSeek-V4
📚 Surveys, Foundations & Position Papers
| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |
|---|---|---|---|---|---|
| GKD | 2023.06 | Google DeepMind (Agarwal et al.) | arXiv 2306.13649 — implemented in TRL GKDTrainer |
GKD: On-Policy Distillation of Language Models — Learning from Self-Generated Mistakes (Seminal · ICLR 2024) | |
| Blog | 2025.10 | Thinking Machines Lab (Kevin Lu et al.) | Blog · tinker-cookbook | Thinking Machines Lab — On-Policy Distillation (blog) | |
| tinker-cookbook | 2025.10 | Thinking Machines Lab | — | Reference impl. of the OPD recipe on the Tinker SDK | |
| revisiting_opd | 2026.03 | CASIA (Fu et al.) | arXiv 2603.25562 | Revisiting OPD: Failure Modes & Simple Fixes | |
| Tencent OPD Survey | 2026.04 | Tencent (Mingyang Song & Mao Zheng) | arXiv 2604.00626 | A Survey of On-Policy Distillation for LLMs | |
| OPD | 2026.04 | Tsinghua THUNLP | arXiv 2604.13016 | Rethinking On-Policy Distillation: Phenomenology, Mechanism & Recipe | |
| Lightning OPD | 2026.04 | Wu, Han, Cai | arXiv 2604.13010 | Lightning OPD: Efficient Post-Training with Offline OPD | |
| OPSD Survey | 2026.05 | Academic | arXiv 2605.18141 | A Brief Overview: On-Policy Self-Distillation in LLMs | |
| Many Faces of OPD | 2026.05 | UIUC (Ge Liu's ULab) | arXiv 2605.11182 | The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes — diagnoses when OPD/OPSD succeed or fail (distribution mismatch, optimization instability, PI-free policy limits) & proposes fixes; companion to Revisiting OPD & THUNLP Rethinking | |
| Blog | 2026.06 | Li Jiang | Blog · arXiv 2606.08432 | On-Policy Distillation: Promise, Pitfalls, and Prospects — reflection on OPD's promise, three failure mechanisms (local teacher noise, coverage decay, myopic gradients) & prospects; companion to TRD | |
| Formula-Driven Survey | 2026.06 | Tsinghua (Bowen Zhang) | arXiv 2606.22793 | A Formula-Driven Survey and Research Agenda for OPD — organises OPD as a feedback-to-update path rather than by KL direction; seven formula-derived variables + an explicit evidence-tier table (E0 papers → E3 own hypotheses) | |
| NAIL | 2026.06 | Columbia (Sriraman, Liu, Hsu, Block) | arXiv 2606.30923 | Behavior Cloning is Not All You Need: The Optimality of OPD for Noisy Expert Feedback — proves offline BC needs samples exponential in horizon under a noisy expert (and that this is necessary for any offline IL), while OPD is polynomial; proposes NAIL | |
| Online IL Realizability | 2026.06 | Tsinghua / CMU / Berkeley / Harvard | arXiv 2606.30445 | When Does Online Imitation Learning Help in LLM Post-Training? — challenges the error-accumulation story: under realizability OPD buys nothing over SFT; non-realizability is the real source of online gains | |
| Demystifying OPD | 2026.07 | CUHK / Tencent AI Lab | arXiv 2607.13399 | Demystifying OPD: Roles, Pathologies, and Regulations — OPD as exploration catalyst, not ceiling-raiser; diagnoses Student–Teacher Mismatch (the strongest teacher can be the worst) + Length Exploitation, and fixes both with clipping / log-compression | |
| Unmasking OPD | 2026.05 | Apple | arXiv 2605.10889 | Unmasking OPD: Where It Helps, Where It Hurts, and Why — training-free diagnostic that derives an ideal per-token gradient from empirical success probabilities and scores GKD / MiniLLM / Dr.GRPO objectives by cosine alignment to it. Finds guidance is far better aligned on incorrect rollouts than correct ones | |
| Decoupling KL & Trajectories | 2026.05 | EIT Ningbo / HK PolyU / HKUST / SJTU | arXiv 2605.16826 | A Unified Perspective for SFT, DAgger, Offline RL, and OPD — decomposes distillation along prefix source × KL direction, showing the four cells are exactly SFT / on-policy SFT / offline-RL distillation / OPD, and proves student-prefix + reverse KL equals a dense-reward REINFORCE objective | |
| States, Not Tokens | 2026.05 | Independent (Dong Nie) | arXiv 2605.22731 | Post-Training is About States, Not Tokens — recasts SFT/RL/OPD along state source × signal source; shows continuation-OPD from a deliberately degraded teacher still lifts the student past that teacher on all three axes | |
| Geometry of OPD | 2026.06 | HKUST / Brown / ZJU / HK PolyU / USTC / BUPT | arXiv 2606.07082 | On the Geometry of On-Policy Distillation — locates OPD between SFT and RLVR in parameter space (update sparsity 51.6% vs 8.1% / 77.2%) and identifies early subspace locking into a persistent low-rank update channel | |
| Dense Supervision, Sparse Updates | 2026.06 | Nanjing Univ. / Amap Alibaba | arXiv 2606.13657 | On the Sparsity and Geometry of OPD — OPD updates are tiny (0.036–0.142% relative Frobenius norm) and 67–90% coordinate-sparse, concentrated in FFN; training only the discovered sparse mask nearly recovers full OPD | |
| Rock Tokens | 2026.05 | UMBC / Case Western / ASU / VU Amsterdam | arXiv 2605.09253 | Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in OPD — tokens keeping high KL after saturation are mostly structural scaffolding whose gradients Adam neutralizes, and are causally near-irrelevant to accuracy | |
| OPSD Compresses RLVR | 2026.05 | Yonsei | arXiv 2605.06188 | OPSD Compresses What RLVR Teaches — separating correct-only from incorrect-only rollouts shows OPSD mainly compresses already-correct traces (−29.22% length at −1.03 pp accuracy, Qwen3-8B) rather than repairing failures; proposes it as a post-RL compaction stage | |
| Extrapolation Cliff | 2026.05 | NTU Singapore | arXiv 2605.08737 | The Extrapolation Cliff in OPD of Near-Deterministic Structured Outputs — derives a closed-form clip-safety threshold λ* beyond which reward-extrapolated OPD collapses output validity; the quantitative bound on G-OPD/ExOPD-style λ>1 | |
| Outcome-Confounded Supervision | 2026.07 | CUHK Shenzhen (Guoqing Ma) | arXiv 2607.23731 | Outcome-Confounded Local Supervision in OPD — the standard local read of teacher/student token-level agreement vs. disagreement is confounded by the completed trajectory's outcome: "agreement-on-failure" tokens carry ~68% of token mass in math reasoning. Three candidate fixes were tested and none consistently resolves it; the paper is explicit that "Our contribution is therefore diagnostic rather than a new training method" | |
| OP²SD | 2026.08 | MBZUAI / Nagoya / RIKEN AIP | arXiv 2608.09228 | Privileged Solutions or Context-Induced Teacher Behavior? Dissecting OPSD — swaps the paired reference solution for one from an unrelated problem while holding student rollout, teacher and objective fixed; OP²SD stays competitive with standard OPSD, implying the gains come substantially from context-induced teacher behaviour rather than the privileged content itself. The sharpest ablation yet against the "learning from privileged information" story | |
| Privileged Likelihood ≠ Value | 2026.08 | Salesforce AI Research | arXiv 2608.09263 | Privileged Likelihood Is Not Automatically Value — three checks separating "does the score track better actions", "does feedback construction change what is compared", and "what does the loss actually reinforce". Under a common additive construction, hindsight-feedback scoring is near-random (AUC 0.505) and an outcome-only baseline beats every token-score variant (64.2% vs 24.2–33.9%, 20B model, AIME 2025) | |
| Test-Time Scaling Lens | 2026.08 | HKBU / SJTU / Shanghai Innovation Institute / Ant Group | arXiv 2608.11829 | Towards Understanding OPD through the Lens of Test-Time Scaling — reads OPD with pass@K / avg@K rather than pass@1: OPD narrows the sampling budget needed to reach a given answer instead of raising the reasoning ceiling. The pass@K advantage returns to the pre-OPD base model as K grows, and more previously-solvable problems become unsolvable than the reverse. Coins "illusory distillation" | |
| Every Coin Has Two Sides | 2026.08 | USTC / Peking / iQuest Research / MBZUAI / Zhejiang | arXiv 2608.16647 | On the Dual Nature of Generalization in OPD — systematic study across distribution shift, cross-domain transfer and multi-teacher setups; finds OPD "transfers a teacher's reasoning behavior rather than its answers to particular problems". Same-origin teacher/student pairs generalise across language and domain; cross-origin pairs mostly fit the training distribution, and multi-teacher combination creates capability trade-offs routing cannot fully control | |
| Rethinking Privileged Information | 2026.08 | FirstPrinciples | arXiv 2608.18271 | Rethinking Privileged Information in OPSD — on Qwen3 (1.7B–8B) across science/math, the correct privileged reference gives no consistent benefit: students improve without correct references, unrelated-problem solutions sometimes outperform correct ones, and student predictions track the base model more closely than the reference. Companion negative result to OP²SD and to Rethinking OPSD for Thinking Models | |
| Fin-SD ablations | 2026.09 | CTGT (Yu, Gorlla) | Blog | Privileged Information in Self-Distillation: An Empirical Study — ablates failure-point hint self-distillation (Fin-SD) on finance reasoning with GPT-OSS-20B. Replacing the entire hint with a content-free placebo ("Your mistake is conceptual") matches the full hint (72.52% vs 72.19%, base 69.16%); a same-family hint writer ~6× the student's size adds nothing (both 72.19%); the hint variants match full-solution OPSD (72.19% vs 70.42%) with 9.5% fewer eval tokens. Zero-content version of the unrelated-solution control in OP²SD and Rethinking Privileged Information |
📋 Click to view technical details
| Resource | Loss / Divergence | Data | Teacher Access | Granularity | Notes |
|---|---|---|---|---|---|
| GKD (Agarwal) | Generalised JSD (FKL/RKL configurable) | Mixed (λ interpolates teacher↔student) |
White-box | Token | The seminal paper that named OPD; introduced student-self-rollout supervision. |
| Thinking Machines blog | Reverse KL (student‖teacher) | Student rollouts | White-box | Token | "Swap KL ref model for stronger teacher" recipe; one-line addition to RL trainer. Replicates Qwen3 result at ~1/10 RL cost. |
| Revisiting OPD | Truncated reverse KL + top-p sampling + special-token masking | Student | White-box | Token (filtered) | Diagnoses 3 failure modes: imbalanced one-token signal, unreliable prefix guidance, tokenizer mismatch. |
| Tencent OPD Survey | (survey) | (survey) | (survey) | (survey) | Catalogues 50+ methods; useful as a reference index. |
| THUNLP Rethinking OPD | Reverse KL with progressive top-K alignment | Student | White-box | Token | Identifies two success conditions: compatible thinking patterns + genuinely new teacher capability. Recipe = off-policy cold-start + teacher-aligned prompt selection. |
| Lightning OPD | Cached teacher log-probs over SFT rollouts (offline OPD) | Student (cached) | White-box | Token | Introduces "teacher consistency" — same teacher must be used for SFT and OPD or else gradient bias. Eliminates the live teacher server. |
| OPSD Survey | (survey) | (survey) | (survey) | (survey) | Categorize eight designs; useful as a reference index. |
| Many Faces of OPD | Reverse KL (FKL/RKL comparison) | Student (self-sampled trajectories) | White-box / self | Token | Diagnostic study (not a training method). Identifies three failure mechanisms — distribution mismatch, optimization instability, and a PI-free policy limit specific to OPSD — and proposes stop-gradient objectives + stabilized training as fixes. Covers both OPD and OPSD. |
| Formula-Driven Survey | (survey) | (survey) | (survey) | (survey) | Refuses the "KL direction × teacher access" taxonomy; models OPD as feedback→update with seven variables: state distribution, feedback source, comparison support Ω_t (sampled-token / top-k / full-vocab), temporal credit A_t, gate-or-weight w_t, vocabulary routing, update route (direct-loss vs policy-gradient score-function). Carves out OPD-hybrids and OPD-adjacent (RLVR, teacher-forced KD, SFT) as out of scope. Proposes two unimplemented designs (GAE-OPD, CR-OPD), self-labelled as E3 hypotheses. |
| NAIL (Behavior Cloning Is Not All You Need) | Forward-KL and reverse-KL variants, on a rollout distribution distinct from the scored policy | Student | White-box | Token | Theory + method. Noisy-expert model π*_η = (1−η)π* + η·ν. Offline BC needs n ≳ (1−η)^{−(H+2)}log|Π|/ε — exponential in horizon and necessary for any offline IL algorithm, unlike the horizon-free clean-expert result; the OPD variant is polynomial in H. Key design claim: the loss must distinguish the rollout distribution (greedy) from the scored policy (temp 1) — which standard OPD does not. NAIL ≈ BC/OPD at low noise, far better at high noise (GSM8K + modular addition). |
| When Does Online IL Help | Reverse KL + entropy under the student (analysis only) | Student | White-box / self | Sequence (contextual bandit, H=1) | Position/theory paper. Under realizability (student class can represent the expert) SFT on expert samples fully matches expert performance and OPD adds neither accuracy nor speed — replicated on Countdown, GSM8K (Llama-3.2-3B) and DeepScaleR (R1-Distill-Qwen-1.5B). Under misspecification, discrepancy-based (TV/Hellinger) bounds go vacuous, and an information-theoretic lower bound with coverage coefficient C^e_∞ applies even at H=1. Notes strong-to-weak OPD lives in the non-realizable regime — which is why it works. |
| Demystifying OPD | Reverse KL as policy gradient, per-token Δℓ_t = log π_T − log π_θ, PPO-clipped | Student | White-box | Token | Analysis + method. (1) OPD steers the student toward correct paths without expanding the capability ceiling (pass@1024 converges to the base model's); prompt diversity beats per-prompt rollout depth. (2) Student–Teacher Mismatch: Qwen3-4B-GRPO, the strongest teacher, is the worst teacher (student stuck ~2% AIME25); an Informativeness metric I = E[Δℓ̄|r=1] − E[Δℓ̄|r=0] predicts this in advance. (3) Length Exploitation: token-mean advantage lets the student wash out negative advantage by padding or truncate early on a favourable prefix. Fixes = Hard Clipping + Soft Log-Scale Compression, zero extra compute. Regulated OPD with a 1.7B/4B teacher beats prior OPD from a 30B teacher (AIME24 45.2 vs Uni-OPD 35.2 / G-OPD 37.3). |
| Fin-SD ablations (CTGT) | Per-token reverse KL over a 100-token window from the located error (beat forward KL in this short-horizon setup) | Student (on-policy prefix, truncated just before the first incorrect step) | Self (student conditioned on a short hint at the error; hint written by a frozen student copy or a ~6× same-family model, or replaced by a fixed placebo) | Token | Empirical study. Hint content and hint-writer size both ablated with no accuracy effect at 20B; windowing 8192→100 tokens improves accuracy, completion and length; residual style-token drift (more "but / maybe / however") and length growth with more epochs. Stated scope: finance reasoning at ≥ GPT-OSS-20B capability, where the needed concept is already latent. |
📝 Strictness notes (against the strict OPD definition C1: student samples its own trajectories during training + C2: teacher provides supervision on those samples)
- Lightning OPD — ⚠️ partially satisfies C1: teacher log-probs are pre-computed once over SFT rollouts and reused during training; student doesn't actively sample during the OPD step. Authors call this "offline OPD" explicitly. Listed in OPD because the data is past-student-generated rollouts, not teacher-generated.
🔬 OPD with Larger External Teachers — White-Box
White-box methods use teacher logits / log-probabilities to supervise the student on student-generated rollouts. Each entry below has been verified to (a) train on student rollouts and (b) operate at the token level.
Methods that turned out to be RL-style on verification have been moved to OPD-RL Hybrids; off-policy / pure-loss-function / pretraining-side methods are excluded from this list.
| Resource | 🌟 Stars | Date | Org | Paper Link | Title / Notes |
|---|---|---|---|---|---|
LMOps /minillm |
2023.06 | Microsoft / Tsinghua | arXiv 2306.08543 | MiniLLM (ICLR 2024) | |
| distillm | 2024.02 | KAIST / Microsoft | arXiv 2402.03898 | DistiLLM (ICML 2024) | |
google-research /speculative_kd |
2024.10 | UCSB / Google | arXiv 2410.11325 | Speculative KD (ICLR 2025) | |
| distillm-2 | 2025.03 | KAIST / Microsoft | arXiv 2503.07067 | DistiLLM-2 (ICML 2025 Oral) | |
| DSKDv2 | 2025.04 | BJTU | arXiv 2504.11426 | DSKDv2 — cross-tokenizer; supports on-policy mode | |
| Constrained OPD | 2025.09 | Huawei Noah's Ark | arXiv 2509.22921 | Constrained OPD (CMDP) | |
| AdaSwitch | 2025.10 | RUC / Baidu | arXiv 2510.07842 | AdaSwitch (on-/off-policy switching) | |
| Veto | 2026.01 | SNU | arXiv 2601.07155 | Veto (Stable OPD) — ACL 2026 Findings | |
| G-OPD | 2026.02 | RUC / Tencent | arXiv 2602.12125 | G-OPD | |
| Fast OPD | 2026.02 | Industrial | arXiv 2602.15260 | Fast OPD (prefix-truncated) | |
| Entropy-Aware OPD | 2026.03 | KAIST / IBM | arXiv 2603.07079 | Entropy-Aware OPD | |
| REOPOLD | 2026.03 | KAIST / Microsoft | arXiv 2603.11137 | REOPOLD (Relaxed OPD) — code soon | |
| OPSD_OnPolicyDistillation | 2026.03 | arXiv 2603.11178 | PACED — frontier curriculum self-distill | ||
| TSD-KD | 2026.03 | Korea Univ. | arXiv 2603.13260 | TSD-KD — token-selective dual KD (ICLR 2026) | |
| SCOPE | 2026.04 | USTC / Meituan / Fudan | arXiv 2604.10688 | SCOPE — signal-calibrated dual-path | |
| OPSD_OnPolicyDistillation | 2026.04 | Meta / LinkedIn | arXiv 2604.14084 | TIP — Token Importance, shares LinkedIn OPSD repo with PACED | |
| Hybrid-Policy-Distillation | 2026.04 | SJTU / Shanghai Innovation Institute / Tencent | arXiv 2604.20244 | HPD — Hybrid Policy Distillation; LlamaFactory + verl backends (ICML 2026). ⚠️ headline Tables 2–4 use a lightweight approximation of on-policy sampling that avoids full-sequence rollouts; true on-policy results appear only in §5.4/Table 5 | |
| BRTS | 2026.05 | JHU (Patel group) | arXiv 2605.09725 | BRTS — Best-of-N Teacher Rollout Selection; augments student-context OPD with a curated teacher-context branch (correctness-first, then student-alignment) to cut single-rollout teacher variance | |
| FiRe-OPD | 2026.06 | THU / HKUST / Meituan (Li et al.) | arXiv 2606.02684 | FiRe-OPD — Filter, then Reweight; decouples optimization granularity — hard trajectory-level filtering (drop bottom-p% rollouts by teacher log-prob) + soft token-level reweighting (teacher-confidence × student-confusion), arguing soft weighting beats hard token selection (cf. TIP); verl-based, with a multi-teacher math+code variant | |
| trd | 2026.06 | McGill / Mila / UT Austin (Jiang et al.) | arXiv 2606.08432 | TRD — Trajectory-Refined Distillation; diagnoses prefix failure of dense per-token OPD, refines student rollouts at trajectory level before distilling; verl-based, also applies to OPSD | |
| OPRD | 2026.06 | ZJU / Ant Group | arXiv 2606.06021 | OPRD — On-Policy Representation Distillation; first OPD to supervise in hidden-state space (aligns teacher/student representations across layers on student rollouts, bypassing the LM head) rather than logits; built on the THUNLP OPD stack | |
| IW-OPD | 2026.06 | Xidian / Georgia Tech / Amazon AGI SF Lab | arXiv 2606.22600 · project | IW-OPD — On the Position Bias of OPD; supervising only the prefix 30% of tokens matches full-token OPD while suffix-30% barely learns; reweights the OPD advantage by a prefix-importance term derived from a trust-region argument | |
| DEAR | 2026.06 | Meituan LongCat / Nanjing Univ. / TJUNLP | arXiv 2606.22830 | DEAR — Finding the Evidence; argues entropy-selective OPD (cf. TIP) captures only decision tokens and structurally misses low-entropy, high-divergence evidence tokens where the student is confident yet wrong | |
| ReNIO | 2026.06 | ECNU / Shanghai Innovation Institute | arXiv 2606.23104 | ReNIO — Reweighting Negative Trajectory Importance; finds training on incorrect student outputs beats correct-only under both OPD and OPSD, then approximates correctness with a prefix-computable pivotal-token proxy | |
| PG-OPD | 2026.06 | TeleAI / SJTU | arXiv 2606.21994 | Prefix-Guided OPD — Mining Golden Trajectories from Rollouts; probes candidates at a short prefix by teacher–student top-k overlap and only continues the promising ones to full length (up to +4.80 avg, 2.46× wall-clock) | |
| SEAD | 2026.06 | Capital One | arXiv 2606.28562 | SEAD — Competence-Aware OPD; entropy zones the tokens (skip / RKL / FKL), cosine-anneals FKL→RKL, and adds the first prompt-level curriculum for OPD | |
| DOPD | 2026.06 | JD Explore / NUS / MMLab CUHK / PKU | arXiv 2606.30626 | DOPD — Dual On-policy Distillation; names "privilege illusion" (part of a privileged teacher's gap is information asymmetry, not transferable capability) and routes each token among four regimes by a privilege-advantage gap | |
| MOPD | 2026.06 | Xiaomi LLM-Core / PKU / HKU / RUC | arXiv 2606.30406 | MOPD — Multi-Teacher OPD for Capability Integration; the method paper behind MiMo-V2-Flash's MOPD stage. Shows same-origin teachers are essential — a stronger but distributionally distant teacher diverges outright | |
| Blockwise Drift Gating | 2026.06 | Independent (Zheng & Jiang) | arXiv 2606.24084 | Blockwise Policy-Drift Gating — student-only old↔current drift gate for OPD under rollout reuse (author-labelled preliminary study) | |
| Direct-OPD | 2026.07 | Tsinghua AIR SIA-Lab / ByteDance Seed / PKU | arXiv 2607.05394 · project | Direct-OPD — Weak-to-Strong Generalization; distills the teacher's implicit reward (post-RL minus pre-RL log-ratio) instead of its distribution, so a weaker post-RL teacher can improve a stronger student. Qwen3-1.7B AIME24 48.3→58.3 in ~4 h on 8×A100 | |
| OPD² | 2026.07 | NAVER AI Lab | arXiv 2607.15161 | On-Policy Delta Distillation; replaces the teacher–student log-ratio reward with a delta signal (teacher minus its own base model), plus a sign-agreement gate. Qwen3-1.7B math avg 34.8→54.6 | |
| TOP-D | 2026.07 | HKUST(GZ) / Microsoft | arXiv 2607.04751 | Trust Region Policy Distillation; interpolates teacher and student in probability space so the token reward is lower-bounded and gradient variance is provably bounded. AIME24 avg@32 50.42 vs 24.58 for standard OPD | |
| COPD | 2026.07 | SJTU / Alibaba Qwen | arXiv 2607.19046 | Contrastive OPD; the teacher re-scores the student's tokens twice — under a "light-thinking" and a "heavy-thinking" prefix — and the contrast becomes the advantage. +2.1 pp with 57% shorter responses; ships a COPSD self-distillation variant | |
| TOPD | 2026.07 | CASIA / UCAS | arXiv 2607.16872 | Trace-Based OPD for masked diffusion LMs; supervises only trace-aligned denoising decisions (commitments surviving into the final answer), avoiding the backward-reconstruction states random-mask supervision creates. MATH500 +5.7 with 4× fewer rollout rounds | |
| AOPD | 2026.05 | HUST / PKU / Meituan | arXiv 2605.06387 | Asymmetric OPD; splits the OPD gradient into positive (exploit) and non-positive (imitate) token regions, replacing the noisy policy gradient with truncated forward KL only on the latter. Preserves math accuracy after continual tool-use training (+0.61 vs −7.49 for OPD) | |
| SimCT | 2026.05 | Tencent Hunyuan / USTC / Shanghai Innovation Institute | arXiv 2605.07711 | SimCT — Recovering Lost Supervision for Cross-Tokenizer OPD; builds a common supervision space of shared-vocab tokens ∪ minimal aligned units (finest jointly-tokenizable spans) so the unchanged OPD loss applies across tokenizers | |
| Prune-OPD | 2026.05 | HKUST(GZ) / MBZUAI / UC Merced / SYSU | arXiv 2605.07804 | Prune-OPD — monitors per-position teacher/student top-k overlap, attenuates rewards on prefix drift and truncates once supervision is locally unexploitable; −37.6–68.0% training time. The mechanism KAT-Coder-V2.5 credits for its drift-aware truncation | |
| ESR | 2026.05 | UCLA / BIGAI | arXiv 2605.27028 | Less is More: Early Stopping Rollout for OPD; names Off-Policy Teacher Decay — conditioning the teacher on the student’s drifting prefix collapses its late-token distribution toward student-level autocomplete. Beats full-rollout OPD in every cell at up to 24× lower cost | |
| ADWIN | 2026.05 | PKU / Tencent | arXiv 2605.28396 | ADWIN — Adaptive Windows for Horizon-Aware OPD; truncates by a prefix-admissibility criterion based on gradient-cosine alignment between short-prefix and full-rollout updates, audited by delayed full-rollout probes. 59.3→60.9 avg at ~3.4× fewer FLOPs single-task (the 4.1× figure is from the separate strong-to-weak setting) | |
| Truncated OPD | 2026.05 | CASIA / Meituan | arXiv 2605.31490 | Are Full Rollouts Necessary for OPD? — proves sequence-level OPD accumulates noisy future signal at O(T³) MSE while token-level does not; truncating to 10% of the horizon matches full-rollout OPD while cutting training time 82% | |
| Prefix Teach, Suffix Fade | 2026.05 | ZJU / Meituan LongCat / Jilin | arXiv 2605.13643 | Local Teachability Collapse in Strong-to-Weak OPD; cuts dense supervision at a BIC-detected change-point in the teacher’s top-1/top-2 margin over the student’s reachable candidates. Avg 36.8→40.1 | |
| KAT | 2026.06 | HKUST(GZ) / HKUST / HK PolyU / EIT | arXiv 2606.09471 | Escaping the KL Agreement Trap in OPD — low KL is not evidence of health: the student drifts into a corrupted prefix and the teacher locally agrees, silencing the corrective signal. Sliding-window detection cuts rollout length 59.7% at 2.4× speedup | |
| TA-OPD | 2026.05 | HK PolyU / InfiX.ai | arXiv 2605.26844 | Not All Disagreement Is Learnable — Token Teachability in OPD; supervises only a support-aligned "teachable" subset, matching or beating full-token OPD with 5–10% of tokens | |
| TRB | 2026.05 | T-Tech | arXiv 2605.31159 | Trust-Region Behavior Blending; leaves the OPD loss untouched and instead controls the rollout behaviour policy μ ∝ π_S^(1−β)·π_T^β under a trust-region constraint, annealing back to pure on-policy sampling | |
| TrOPD | 2026.06 | Samsung Research Beijing / Oxford / PKU | arXiv 2606.01249 | Trust Region On-Policy Distillation; partitions student tokens into a teacher-verifiable trust region (reverse KL) vs outliers (top-k forward KL), plus annealed off-policy guidance. AIME25 44.06 vs OPD 40.72 | |
| Near-Future Guidance | 2026.06 | UMBC | arXiv 2606.00305 | Bridging Reasoning Trajectories in OPD via Near-Future Guidance; shows per-token KL is a weak proxy for real trajectory divergence (Pearson r=0.126) and adds an optimal-transport-aligned target spanning several future positions | |
| PowerOPD | 2026.06 | EIT Ningbo / HK PolyU / SJTU / Waterloo | arXiv 2606.17199 | Stabilizing OPD with Bounded Power Transformation; the vanilla reward log(π_T/π_θ) is unbounded and drives instability, so it is replaced by a sign-consistent Box-Cox transform bounded in [−1,1]. +6.37 Avg@8 at 59% less wall-clock | |
| NPD | 2026.05 | Huawei / Tianjin Univ. | arXiv 2605.05940 | Near-Policy Distillation; decouples generation from updates so teacher top-k logits can be computed by packed parallel prefill — 8.1× speedup, at the cost of deliberate policy lag | |
| RWOPD | 2026.05 | NUS | arXiv 2605.13501 | Reward-Weighted OPD with an Open Property-Equivalence Verifier; forward KL on rollouts passing a SymbiYosys+Z3 equivalence check, for NL→SystemVerilog Assertions. Beats its own 14B teacher and 671B general baselines | |
| ProteinOPD | 2026.05 | Tsinghua / IDEA / HKUST(GZ) / NTU | arXiv 2605.10189 | ProteinOPD — multi-teacher OPD for protein language models; distills a normalized product-of-experts consensus over preference-specific teachers via token-level generalized JSD. PPL −83.7%, thermostability +54.2% | |
| MAD-OPD | 2026.05 | HUST / Alibaba | arXiv 2605.01347 | Breaking the Ceiling in OPD via Multi-Agent Debate; two larger teachers (K=2) debate across rounds to form the token-level target, with task-adaptive divergence. A 4B student beats its own 14B teacher on LiveCodeBench-v6 | |
| W2S-OPD | 2026.07 | UMD / Microsoft Research / MBZUAI | arXiv 2607.26246 · project | W2S-OPD — Weak-to-Strong On-Policy Distillation; synthesises a proxy teacher distribution softmax(z_base + α(z⁺ − z⁻)) by re-anchoring a weak contrast pair's logit delta onto the student's own base model, then distils it with per-token reverse KL. Three interchangeable contrast sources: post-RL vs pre-RL expert, larger vs smaller base model, correct vs wrong hints |
|
| Cross-Tokenizer OPD | 2026.07 | UCAS / KwaiKAT Team / Zhejiang Univ. | arXiv 2607.22334 | Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization — maps the teacher's full next-token distribution into the student's vocabulary by marginalising over byte prefixes, so the unchanged β-divergence (FKL/JSD/RKL) applies across mismatched tokenizers with no hand-built alignment table. Third cross-tokenizer construction alongside SimCT and SimpleOPD | |
| Relay-OPD | 2026.07 | Zhejiang Univ. / Yuvion Team, Alibaba | arXiv 2607.26057 | Pass the Baton: Trajectory-Relayed On-Policy Distillation — attacks prefix failure via a teacher–student continuation asymmetry on failed prefixes (the teacher redirects, the student ploughs on) used as a label-free handoff trigger: the teacher briefly writes a "teacher leg", the student resumes, and the relayed trajectory is distilled under a limited relay budget concentrating intervention early. +5.73% over OPD, +1.49% over FastOPD. Same diagnosis as TRD, different remedy (in-trajectory relay vs. post-hoc refinement) | |
| TU-OPD | 2026.07 | Moore Threads AI | arXiv 2607.27770 | Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold — multi-teacher OPD that quality-gates each teacher per rollout: a weight α_i(r) selects which specialist's distribution supervises which student sample, plus a residual term preserving specialist behaviour instead of averaging it away. Verifier score only weights/selects teachers; it never enters the objective | |
| Flux-OPD | 2026.07 | PKU / Kling Team / Tsinghua / SJTU / Zhongguancun Academy | arXiv 2607.28022 | Flux-OPD — On-Policy Distillation with Evolving Contexts — for open-ended domains where rewards are unverifiable: a context expressing the desired preference is fed to the teacher and the difference between context-conditioned and context-free teacher distributions becomes the signal, with the context itself evolving across iterations. Sibling of OPD²'s delta idea, but contrasting the same teacher under two conditionings — two loaded models instead of three | |
| Adaptive FastOPD | 2026.07 | USTC / Shanghai AI Lab / SJTU | arXiv 2607.29494 | Adaptive FastOPD — Progress-Aware Rollout Horizon Expansion — makes the prefix-truncation family adaptive: rather than a fixed truncation length the rollout horizon expands as measured reasoning progress warrants, recovering full-rollout quality at truncated-rollout cost. Direct successor to Fast OPD; belongs with Prune-OPD / ESR / ADWIN / Truncated OPD | |
| OPTD | 2026.08 | HKUST / NPU / HK PolyU | arXiv 2608.02942 | OPTD — On-Policy Transition Distillation for Few-Step Diffusion LMs — samples partial states from the student's own denoising trajectory and matches transitions against a frozen teacher, with a set-bottleneck target and consistency-guided adaptive compression. ⚠️ Same dLLM caveat as TOPD: "trajectory" is a denoising path over masked positions, not left-to-right generation | |
| SA-OPD | 2026.08 | Zhejiang Univ. / ByteDance / Shanghai AI Lab | arXiv 2608.03632 | When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation — a teacher's confident tokens are not always input-grounded; some reflect priors rather than the actual problem. Filters those spurious tokens out of the OPD signal by an input-groundedness gap test (Δ_IG threshold) before applying reverse KL. Only the LLM track is in scope; the paper carries a second VLM track | |
| SPOT (sparse probing) | 2026.08 | CUHK-Shenzhen / ECNU / Meituan / Beihang / CityU HK | arXiv 2608.04419 | SPOT — Sparse Probing and Outcome Calibration — cuts teacher cost by sparsely probing only informative positions of the student rollout (chosen by teacher entropy / top-k candidate mass) and calibrating the resulting targets against the trajectory outcome, giving a closed-form KL-regularised target across acquisition / exploration / exploitation phases. ⚠️ Acronym collision with the Black-Box SPOT (2603.01683, "Surgical Post-Training") — unrelated; resolve by arXiv ID | |
| OPD² (multilingual) | 2026.08 | NAVER AI Lab | arXiv 2608.05802 | On-Policy Delta Distillation for Multilingual Math Reasoning — carries the OPD² delta signal log π_teacher − log π_base into cross-lingual math (EN/KO/JA), where the teacher's absolute distribution is contaminated by language priors the student already holds. ⚠️ Shares both the OPD² name and the same repo as 2607.15161 — a companion second paper, not a replacement; the existing note about needing the teacher's pre-post-training base checkpoint applies here too |
|
| Simple-OPD | 2026.08 | Tsinghua / HKU / Tencent (LLM Dept.) | arXiv 2608.06802 | Simple-OPD — Demystifying Warm-up for On-Policy Distillation — asks what the SFT warm-up before OPD actually buys and finds a cheap LoRA warm-up suffices, after which plain reverse-KL OPD matches far more elaborate recipes. Speaks directly to the "off-policy cold start" finding in THUNLP's Rethinking OPD | |
| WDL-OPD | 2026.08 | Beihang / China Telecom eSurfing Cloud | arXiv 2608.09447 | WDL-OPD — Weak-Driven OPD via Mixture-Constrained Co-Training — only a designated anchor model rolls out ("Only the anchor generates trajectories: y ∼ π_A(·|x)"); the supervision target is a geometric mixture of the anchor's and an auxiliary model's token distributions, constrained toward a frozen teacher — co-training two students against one teacher rather than distilling each separately | |
| TIDE | 2026.08 | HKU / USTC / CUHK | arXiv 2608.09836 | Mismatch Matters: OPD Beyond Token Agreement — splits teacher–student mismatch into excess mass (student over-commits) and deficit mass (student never proposes what the teacher wants), penalising the first with a Hellinger-shaped term and repairing the second by injecting teacher top-K tokens, arguing plain token agreement hides both failure modes. verl-based | |
| ReOrder-OPD | 2026.08 | Hello Group / Institute of Automation CAS / UCAS / PKU | arXiv 2608.10905 | ReOrder-OPD — Reliability-Aware Prompt Ordering — leaves the OPD loss untouched and reorders the prompt curriculum by how reliable the teacher's supervision is on each prompt; a data-ordering counterpart to SEAD's competence-gated curriculum. The ROUGE proxy orders prompts only — the token-level target uses real teacher logits | |
| REOPD | 2026.08 | Shanghai AI Lab / PKU / SWJTU / USTC / UESTC / RUC | arXiv 2608.11698 | REOPD — Reliability-Adaptive Reward Extrapolation — generalises ExOPD/G-OPD reward extrapolation from a single global λ to a per-token compatibility weight gating a batch-adaptive extrapolation budget, cleanly separating base alignment from the extrapolation residual. Filed here rather than in OPD-RL: the advantage traces entirely through log π_θ / log π_T / log π_ref, and the paper states it "requires no verifier, reward model, value model, or extra rollout beyond standard OPD". ⚠️ Acronym near-collisions with ReOPD (2607.04763) and REOPOLD (2603.11137) |
|
| CROP | 2026.08 | University of Hong Kong | arXiv 2608.13387 | CROP — Task Relevance via Counterfactuals for Selective OPD — scores each token's task relevance by counterfactual intervention (would the outcome change if this token changed) and distils only the tokens that matter. Adds to the token-selection debate alongside TIP / DEAR / TA-OPD / R2-OPD. Mechanism is a hard token mask over the unchanged clipped OPD surrogate — no reward or verifier signal, hence White-Box rather than OPD-RL | |
| SimpleOPD | 2026.08 | Shanghai AI Laboratory (SU-01 Team) | arXiv 2608.14277 | SimpleOPD — Tokenizer-Agnostic OPD for Long-Context Reasoning — matches token log-probabilities across differing vocabularies with no alignment table, distilling a long-context teacher into a short-context student. Third cross-tokenizer entry alongside SimCT and 2607.22334, each with a different alignment construction | |
| DUET | 2026.07 | WeChat, Tencent Inc. | arXiv 2608.14644 | DUET — Dual-Teacher OPD via Same-Weight Disagreement — pairs two teachers with identical weights, one shown the prohibition and one not, and treats their token-level disagreement as an isolated, noise-free supervision signal for runtime-injected policy compliance. 72.3–85.2% violation compliance at 88–93% retained utility. Same same-weight-different-context contrast as COPD and Flux-OPD, applied to safety | |
| AED | 2026.08 | Sun Yat-sen Univ. / Hong Kong Metropolitan Univ. | arXiv 2608.14685 | Rethinking Reverse KL as Adaptive Entropy Distillation — decomposes the reverse-KL OPD objective into a teacher-fitting term and a student-entropy term, then uses the teacher's own per-token entropy to calibrate imitation strength: high-entropy (uncertain-teacher) tokens get less forced imitation, low-entropy tokens more — with no explicit forward-KL branch. Full-vocab and teacher-top-k variants; MATH-500 Avg@8 14.63%→47.80% over vanilla RKL in one reported config | |
| TA-OPD (tail-aware) | 2026.08 | SUSTech (Statistics & Data Science) | arXiv 2608.14728 | Tail-Aware Top-k On-Policy Distillation — top-k teacher truncation silently discards the tail probability mass; TA-OPD keeps the aggregated tail as an explicit extra distillation target, in three variants (normalised top-k, TA-OPD, sample-corrected TA-OPD). verl/vLLM on DAPO-MATH-17K. ⚠️ Acronym collision with the listed TA-OPD (2605.26844, token Teachability, repo wyy-code/TA-OPD) — different paper, different repo, different mechanism |
|
| Open-MOPD | 2026.08 | SIA-Lab (Tsinghua AIR & ByteDance Seed) / Tsinghua | arXiv 2608.19098 | Open-MOPD — Diagnosing and Fixing Capability Imbalance in Multi-Teacher OPD — the open reproduction of the Xiaomi MOPD recipe: diagnoses naive multi-teacher merging (math / code / instruction-following) as an optimization-budget imbalance rather than a capability conflict, and rebalances it — integration gap 3.50 → 0.31 points, recovering 83.4% of the achievable maximum on SmolLM3-3B. Ships 5 checkpoints + 3 domain RL teachers + data + full pipeline, making it the most reproducible multi-teacher OPD artifact in the list | |
| R2-OPD | 2026.08 | HKUST(GZ) / Tsinghua / Zhejiang / Univ. of Alberta | arXiv 2608.19408 | Beyond Imitation: Filtering OPD by Reasoning Progress — argues teacher likelihood and actual reasoning progress can point in opposite directions, and masks out tokens where the teacher-imitation ranking disagrees with an independently estimated progress reward, keeping only supervision that moves the solution forward. A non-entropy, non-divergence selection criterion for the token-selection debate |
📋 Click to view technical details
| Method | Loss / Divergence | Data | Granularity | Domain | Notes |
|---|---|---|---|---|---|
| MiniLLM | Reverse KL via policy gradient | Student | Sequence (PG) | General | The seminal "OPD" recipe by Yuxian Gu et al.; predates GKD by days. Mode-seeking. |
| DistiLLM | Skewed-KL (mix of FKL/RKL) | Mixed (adaptive off→on, with student samples) | Token | General | Skew parameter α interpolates between FKL and RKL; importance-reweighted student samples. |
| Speculative KD (Xu) | Interleaved propose-and-correct (gated KL) | Student-proposed, teacher-corrected | Token | General | Bridges teacher-student gap via interleaved sampling. |
| DistiLLM-2 | Contrastive: Skew-FKL on teacher data + Skew-RKL on student data | Mixed | Token | General | Asymmetric losses on each data source; ICML 2025 oral. |
| DSKDv2 | KL in dual aligned space; explicit on-policy mode | Student | Token | Cross-tokenizer | Cross-vocabulary distillation; supports both on/off-policy. |
| Constrained OPD | KL-constrained CMDP | Student | Token | General | Hard KL constraint instead of soft penalty. Borderline OPD-RL. |
| AdaSwitch | Adaptive on/off-policy switching | Mixed | Token | General | Switches between teacher-data and student-rollout based on divergence threshold. |
| Veto | Logit-space geometric bridge with adaptive gradient veto | Student | Token | General | Adaptive Target Reformulation. |
| G-OPD / ExOPD | Reverse KL + scaled reward extrapolation | Student | Token | General | Generalises OPD as KL-constrained RL; allows reward scale > 1 to "exceed" the teacher. |
| Fast OPD | Prefix-truncated distillation reducing FLOPs | Student | Token (truncated) | Reasoning | 2× to 47× speedup via reasoning-prefix truncation. |
| Entropy-Aware OPD | Switch between FKL and RKL based on teacher entropy | Student | Token | Reasoning | When teacher entropy high → FKL; low → RKL. |
| REOPOLD | Mixture-based reward clipping + entropy-based dynamic sampling | Student | Token | Reasoning | "Relaxed OPD"; views OPD as policy optimisation with teacher-student log-ratio reward. |
| PACED | Frontier curriculum at student competence boundary | Student | Token | General | Self-distill style (privileged-context / earlier-checkpoint); difficulty weighting w(p)=p(1−p). |
| TSD-KD | Indirect (student-propose / teacher re-rank) + direct selective logit KD | Mixed | Token (selected) | General | Hybrid; partial OPD + partial preference. |
| SCOPE | Teacher-PPL-weighted KL on incorrect rollouts; student-PPL-weighted MLE on correct | Student | Token | Reasoning | Signal-Calibrated OPD with Dual-Path Adaptive Weighting; verifier-routing. |
| TIP | Top-50% high-entropy student tokens carry the OPD signal | Student (selected) | Token (filtered) | Reasoning | ~47% memory savings; only entropy-high student tokens trained. |
| HPD | Reweighted log-likelihood unifying FKL + RKL | Mixed (off-policy + lightweight approximate on-policy sampling) | Token | General | Unifies KD as token-level reweighted likelihood; lightweight on-policy sampling preserves training efficiency. |
| TRD | Trajectory-level refinement of student rollouts, then distillation | Student (refined) | Trajectory → Token | Reasoning | Argues dense per-token supervision causes prefix failure; revises problematic student predictions at the trajectory level before distilling. Generalises to on-policy self-distillation. |
| BRTS | Token-level KL on student rollouts + auxiliary teacher-context loss | Student + Best-of-N-selected teacher rollouts | Token | Reasoning (AIME/AMC) | Selection waterfall = correctness → student-alignment → ground-truth-guided recovery; the curated teacher trajectory replaces a single high-variance teacher rollout. |
| OPRD | Layer-wise representation alignment (deterministic; avoids Monte-Carlo KL variance over large vocab) | Student | Hidden-state (per-layer) | Reasoning (competition math) | Lifts distillation from output space into hidden-state space — "bypasses the LM head entirely". Teacher access is white-box (representations rather than logits): the first feature-based OPD. |
| FiRe-OPD | RKL with hard trajectory filtering + soft token reweighting | Student (filtered) | Trajectory (hard) → Token (soft) | Reasoning + Code | Decouples granularity: hard filtering wins at the trajectory level, soft weighting beats hard selection at the token level. Weight w_t=(1+α·c^T_t)(1+β·c^S_t) from teacher confidence × student confusion, per-trajectory normalised. +6.25 AIME24 (Qwen3-4B, 30B teacher); multi-teacher math+code variant. |
| IW-OPD | Per-token OPD advantage × detached, normalised prefix-importance weight | Student | Token (position-reweighted) | Reasoning + code | Diagnoses a position bias: student rollouts drift out of the teacher's distribution as they lengthen, so late tokens carry degraded supervision. Derived as the closed-form optimum of min_q D_KL(q‖π_T) s.t. D_KL(q‖π_θ) ≤ ρ, changed of measure back to π_θ. Stabilised with log-space scaling, unsigned accumulation, within-sample normalisation. No extra teacher forward passes. +6.9 AIME-2025 at step 10; gains grow as the student shrinks (+1.0/+1.2/+1.9 for 4B/1.7B/0.6B) and as the compression ratio rises (+4.0% at 1.0× → +14.9% at 6.7×). Composable with ExOPD. |
| DEAR | Teacher–student log-prob gap advantage on a selected token subset (D ∪ E, ~36% of tokens) | Student (selected) | Token (filtered) | Reasoning + code | Direct methodological rebuttal to TIP. Entropy selectors find decisions (where to branch) but the substantive knowledge sits in evidence tokens — low-entropy, high-divergence positions where the student is confident yet wrong, structurally unreachable by any entropy criterion. Stage 2 scores non-decision tokens by hidden-state cosine similarity to decision anchors × normalised divergence. Gradient-mass coverage: 39.1% (decision-only) / 35.9% (random) / 75.8% (DEAR). Up to +2.5 pp competition math, +5.7 pp code. |
| ReNIO | Standard token-level OPD divergence × a per-sample weight from clipped "pivotal tokens" | Student (reweighted) | Sample-level weight on token loss | Math + code | Controlled filtering shows incorrect-only training beats correct-only under both OPD (+2.59) and OPSD (+2.50), and yields longer responses with more reflection markers. Because correctness labels need full answer-bearing rollouts (defeating OPD's short-prefix cost advantage), the weight is built from prefix-computable log-ratios instead. Evaluated in both external-teacher OPD and teacher-free OPSD (teacher = the model's own initial parameters) modes; the OPSD tables are the stronger half. |
| PG-OPD | Unchanged reverse KL; the contribution is rollout budgeting | Student (prefix-screened) | Token | Math reasoning | All K candidates decode to a fixed prefix; teacher–student top-k overlap over the first R probe tokens scores each; only high scorers (plus a guaranteed per-prompt best) continue to full length, while prefix tokens of all candidates still supervise. Up to +4.80 avg over OPD (59.40→64.20) at 2.46× wall-clock; beats PRUNE-OPD. |
| SEAD | Zone-dependent: skip (both low-entropy) / RKL (teacher confident, student uncertain) / FKL (teacher uncertain) | Student | Token (zoned) + prompt (curriculum) | Math reasoning | Three scales of competence-adaptivity. Token: joint top-k entropy assigns ~50% of positions zero gradient, ~40% RKL to sharpen, ~10% FKL to preserve multi-path diversity. Temporal: α cosine-anneals 0.8→0.0, a continuous exploration→refinement transition. Prompt: a competence-gated curriculum admits prompts as the student's measured pass rate allows. 64.0 avg vs 59.2 vanilla OPD; a 2³ factorial shows curriculum alone is the single strongest factor (+4.20) while zones + annealing are super-additive. |
| DOPD | Four-way routed: Top-K RKL→teacher / stop-grad self-anchor / full-vocab JS→teacher / RKL→privileged student | Student (non-privileged rollout) | Token (routed) | Reasoning + VLM | The privileged-teacher trick raises the ceiling but conflates two gaps: the transferable capability gap the student is meant to close, and the information-asymmetry gap it can only mimic. Uniform distillation therefore teaches privileged shortcuts and collapses entropy. A per-token gap A = |log π_T − log π_S| measured under identical privileged conditions selects the regime. Qwen3-8B→1.7B avg 51.4 vs 43.9 vanilla OPD, recovering 89.8% of the teacher–student gap. |
| MOPD (Xiaomi) | Per-token reverse KL; PG form with clipped  = sg[log π_teacher − log π_student], or top-k (k=64) with a bias-correction term | Student | Token | Multi-domain (Math, IF, SWE, Code, Tool Use) | The canonical inter-stage consolidation recipe: one shared SFT checkpoint → fully parallel per-domain RL experts → student re-initialised from that same SFT checkpoint and distilled from domain-routed teachers on its own rollouts. Teachers run as standalone prefill services so teacher cost hides behind student sampling. Normalised score 0.9373 vs 0.8818 (Mix-RL) / 0.8241 (off-policy finetune) / 0.8574 (param-merge). Same-origin teachers are essential — substituting the stronger but distant Qwen3-235B-A22B collapses it to 0.60 (PG) / −1.19 (top-k) with divergence at step 18. Iter-2 reaches 0.986. |
| Blockwise Drift Gating | Existing OPD loss × detached gate g = exp(−τ|s|) from block-aggregated old↔current drift | Student (rollout-reuse) | Block (64-token / newline span) | Math reasoning | Targets the PPO-style setting where one rollout is reused across epochs, so π_θ drifts from the π_old that generated it. Teacher targets, support and rollout policy are all untouched. LSM+Block64 is the best trained student at 53.3 avg (vs LSM 49.8, base 40.1, teacher 65.2). |
| Direct-OPD | Teacher log-ratio r_t = log π_T − log π_T_ref (the teacher's implicit reward), Rao–Blackwellised over top-k, with adaptive α anchoring to the student's init | Student | Token | Reasoning | Weak-to-strong: run cheap RL on a small model, then transfer what that run learned to a bigger target. Vanilla OPD toward a weak post-RL teacher actively degrades a stronger student (R1-Distill-7B 56.7→~50); reading the pre-RL→post-RL shift instead improves it. The objective is itself KL-regularised RL anchored at the student's own initialisation. Qwen3-1.7B AIME24 48.3→58.3 (~4 h, 8×A100, vs Polaris direct RL at 32×A100 for a week); shifts compose (48.3→58.3→63.8). Transfer works without rising teacher–student top-k overlap, so it is not progressive imitation. |
| OPD² (Delta Distillation) | Delta reward R^Δ = log π*(y_t) − log π*_base(y_t), both advantages top-k-centred, with a sign-agreement gate | Student | Token | Math + science + code | Isolates what post-training added to the teacher rather than the teacher's absolute distribution, so the student is not pulled toward capabilities it already shares with the base model. The sign-agreement gate zeroes updates where the delta and ordinary OPD advantages disagree, fixing the convergence point. Implemented on TRL GRPOTrainer, 100 steps. Qwen3-1.7B math avg 34.8→54.6 (vs 51.0 OPD, 51.4 ExOPD); Qwen3-8B AIME24 76.2. |
| TOP-D | Probability-space teacher–student interpolation collapsing to r̃ = log(αρ + 1 − α); GRPO-style clipped surrogate with token-level normalisation | Student (rollouts reused across mini-batch epochs) | Token | Math reasoning | The smoothed reward is strictly lower-bounded, so gradient variance is bounded (Thm 4.2) instead of exploding as π*→0 — at zero extra compute over standard OPD. Comes with a global convergence bound and a monotonic-improvement bound. Qwen3-8B-Base ← Qwen3-30B-A3B: AIME24 avg@32 50.42 vs 24.58 standard OPD, AIME25 +10.73, AIME26 +18.64; vs GRPO 30.10 / DAPO 32.92. |
| COPD | Log-likelihood contrast A_t = ℓ^LT − ℓ^HT between two teacher scorings of the same student tokens, clipped and used as a PPO-style advantage | Student | Token | Multimodal reasoning | Instead of matching a distribution, it asks the teacher which of two reasoning modes the token belongs to. No accuracy reward, no verifier, no length penalty. vs ExOPD on Qwen3-VL-8B→2B: +2.1 pp with 57.0% shorter responses (Acc@1K 26.7→64.4) at 60 GPU-h vs 132 (OPD) / 150 (ExOPD). The COPSD variant swaps in a frozen snapshot of the student itself: +3.5 pp, −63.8% length. |
| ShortOPD | Generalized JSD (α=0.5) over top-100 logits + aggregated tail mass | Student (adaptive horizon) | Token | Compression (post-pruning recovery) | Pruning Qwen3-4B-Instruct by 4/36 layers collapses greedy GSM8K 88.1→49.0 but pass@64 recovers to 91.2 — correct trajectories are demoted, not erased, so recovery needs on-policy states and dense targets with the frozen pre-compression parent as teacher (no labels, verifier, or external teacher). At 25% pruning 55–75% of early rollouts end in repetitive suffixes carrying ~35× less signal, so EMAs of repetition / truncation / effective length adapt the per-step horizon. Avg over 8 tasks 5.71→48.46 = 64.5% of the dense teacher, +17.9 over the best off-policy baseline, 8.5 h vs 35.9 h. |
| W2S-OPD | Per-token reverse KL toward a synthesised teacher softmax(z_base + α(z⁺ − z⁻)), estimated on the teacher's top-K support |
Student | Token | Math + code | Weak-to-strong: every supervision source is smaller than the student. Because the weak pair enters only through its logit difference, what the two weak models agree on (their shared limited ability) cancels, and only the direction along which m⁺ improves over m⁻ transfers; re-anchoring that direction on the student's own base keeps the target inside a distribution the student already realises, with α bounding how far it moves. Eq. 2 is an exponential tilt of the student's base, so the proxy teacher is the closed-form argmax of E_q[r] − (1/α)·KL(q‖π_base) — a trust region centred on the student, not on any weak model. Three contrast sources are interchangeable (post-RL vs pre-RL expert · Qwen3-4B vs 0.6B base · one model under correct vs wrong hints), and several deltas sum on the shared anchor for multi-teacher merging. Qwen3-8B math avg 17.0→51.8 (vs 46.5 OPD, 45.7 SFT), surpassing the 4B expert itself (48.8); +6.0 pp from two off-the-shelf base models both weaker than the student; +1.4 pp from a single model under contrastive hints. OOD GPQA-Diamond 38.9→56.5, and IFBench improves where OPD drops below the base. A top-1% highest-Δ token analysis over Schoenfeld episodes separates the sources: the post-RL and hint contrasts reinforce Plan/Monitor, the scale contrast Analyze/Implement. |
| TOPD (masked dLLM) | Token-level reverse KL via a sampled-token score-function estimator, on trace-aligned decisions only | Student (own denoising trajectory) | Token (trace-aligned) | Reasoning (diffusion LMs) | Random-mask supervision — the default in dLLM RL/SFT pipelines — can reveal later answer tokens while hiding earlier reasoning ones, creating backward-reconstruction states the student never visits at inference. TOPD instead rolls out the real low-confidence-remasking decoder and keeps only commitments that survive into the final answer. SDAR-4B-Chat ← TraDo-8B-Instruct: MATH500 70.2→75.9, matching TraceRL-trained TraDo-4B with 4× fewer rollout rounds. Ablations: on-policy 75.2 > off-policy 74.5 > semi-AR SFT 73.5; trace-aligned 75.2 > random-mask 74.3; RKL > JSD > FKL. |
📝 Strictness notes
-
BRTS — ⚠️ Partially dilutes C1: the primary student-context leg is strict OPD (student trains on its own rollouts), but the auxiliary teacher-context branch supervises on teacher-generated (off-policy) trajectories. Listed because the student-context leg is the core objective and the teacher branch only stabilises it.
-
OPRD — ⚠️ Not logit-based: C1 ✓ / C2 ✓ on student rollouts, but supervision is feature/representation-space (hidden states across layers), not next-token logits. Listed in White-Box because teacher access is white-box; flagged here because the "feature" supervision signal sits outside the section's default logit-matching form.
-
Direct-OPD — ⚠️ Two departures from the section default. The supervision is a teacher log-ratio (implicit reward), not the teacher distribution, so it is arguably an OPD/RL hybrid; and the teacher is smaller than the student (weak-to-strong), which inverts the usual strong-to-weak setting. C1 ✓ / C2 ✓ — the teacher is still queried on the student's own visited prefixes.
-
OPD² — ⚠️ Stronger access assumption than ordinary white-box OPD: three models must be loaded (student, teacher, and the teacher's pre-post-training base checkpoint), which is not available for most released teachers.
-
COPD — ⚠️ C2 is satisfied at token granularity (teacher log-probs on student tokens) but the loss is not a KL/distribution-matching objective — the teacher signal is converted into an RL advantage. Listed in White-Box rather than OPD-RL Hybrids because there is no reward model or verifier anywhere in the objective. Its COPSD variant is squarely OPSD.
-
TOP-D — ⚠️ C1 slightly relaxed: the internal trust-region iterations deliberately reuse rollouts across mini-batch epochs, so updates are near-on-policy. The paper's own "w/o off-policy" ablation is the strict variant.
-
Blockwise Drift Gating — ⚠️ Authors describe it as "a preliminary empirical study": one student, one teacher, one dataset, no repeated seeds, and AIME sets with tiny sample counts — a 1.7-point pass@8 delta on 4 benchmarks is plausibly within noise. Also a pure loss-weighting heuristic on an existing OPD loss, not a new supervision mechanism.
-
ShortOPD — teacher is the student's own uncompressed parent, so it sits between the compression slot and OPSD. C1 ✓ / C2 ✓ otherwise textbook.
-
W2S-OPD — ⚠️ Inverts the section default: every supervision source is smaller than the student (weak-to-strong), so "larger external teacher" describes only the synthesised proxy, not any real model on disk. Three frozen models must be loaded (the student's own base as anchor + the contrast pair), though only the post-RL instantiation needs a post-training checkpoint pair; the scale and hint variants use two off-the-shelf base models or a single model under two prompts. C1 ✓ / C2 ✓ — the delta is composed in logit space and supervision remains a full top-K distribution matched by reverse KL, so the objective stays distribution-matching rather than an OPD/RL hybrid.
-
⚠️ Acronym collisions. This field has reused several abbreviations for unrelated work; always resolve by arXiv ID:
- TOPD = Trace-Based OPD for masked dLLMs (2607.16872) · Near-Future-Guidance trajectory OPD (2606.00305) · Truncated OPD (2605.31490).
- MOPD = Multi-Teacher OPD, Xiaomi (2606.30406) · Multi-Rollout OPD, Microsoft/CMU/Purdue (2605.12652) · and generically for multi-teacher consolidation in most 2026 production reports.
- COPSD = Crosslingual OPSD (2605.09548) · Constitutional On-Policy Safe Distillation (2606.03089).
- D-OPSD / d-OPSD / dOPSD = step-distilled image diffusion (2605.05204) · dLLM self-future (2606.18195) · dLLM peek-ahead (2607.04428).
- PBSD = Preference-Based Self-Distillation (2605.05040, OPSD) · Posterior-Bayesian Self-Distillation (2606.09348, Agent) — unrelated mechanisms, same acronym.
- COPD (2607.19046, SJTU/Qwen — contrastive light-vs-heavy-thinking prefixes) and CoPD (2604.27083, JD.COM — co-evolving sibling branches) differ only in capitalisation.
- TrOPD (2606.01249, Samsung/Oxford/PKU) and TOP-D (2607.04751, Microsoft/HKUST-GZ) are different papers with different mechanisms — verified separately.
-
Teacher decay is a recurring, independently-rediscovered phenomenon. Supervision quality degrades as the student's prefix lengthens, named separately as Off-Policy Teacher Decay (ESR), Supervision Fidelity Decay (LGR), local teachability collapse (2605.13643) and depth-inverted discriminability (TurnOPD). KAT adds the sharper warning that low KL is not evidence of health — the teacher may simply be agreeing with a corrupted prefix. The prefix-truncation family (Fast OPD, Prune-OPD, PG-OPD, ESR, ADWIN, Truncated OPD) is the practical response.
-
Whether incorrect or correct rollouts carry the signal is genuinely contested. ReNIO and Apple's diagnostic find supervision is better aligned on incorrect rollouts; Yonsei's compaction study finds OPSD mainly compresses already-correct traces and barely repairs failures. Both were verified by full reads; the disagreement is real, not a listing error.
-
TOPD / diffusion-LM entries — C1 ✓ (the student rolls out its own denoising trajectory with the real inference-time decoder) and C2 ✓ at token level, but "trajectory" means a denoising path over masked positions rather than a left-to-right generation, so per-step semantics differ from the autoregressive entries.
-
REOPD (2608.11698) and CROP (2608.13387) — ⚠️ both were initially filed as OPD-RL hybrids and moved here after a full-text read. Both wear a PPO-shaped surrogate, but in neither case does any signal outside the teacher's own log-probabilities enter the objective: REOPD's advantage traces entirely through
log π_θ / log π_T / log π_refand the paper states it "requires no verifier, reward model, value model, or extra rollout beyond standard OPD"; CROP is a hard token mask over the unchanged clipped OPD surrogate. Same filing logic as G-OPD (2602.12125), whose reward extrapolation also lives here. -
OPTD (2608.02942) — ⚠️ diffusion-LM caveat, as already applied to TOPD (2607.16872): the supervised "trajectory" is a denoising path over masked positions, not left-to-right generation.
-
OPD² multilingual (2608.05802) — ⚠️ shares both the OPD² name and the same GitHub repo (
naver-ai/opd2) as the existing 2607.15161 entry. Companion second paper, not a replacement; the existing note about requiring the teacher's pre-post-training base checkpoint applies here too. -
New acronym collisions this window — SPOT: 2608.04419 (sparse probing, White-Box) vs. the listed 2603.01683 ("Surgical Post-Training", Black-Box). TA-OPD: 2608.14728 (Tail-Aware,
HuipengHuang/TA-OPD) vs. the listed 2605.26844 (token Teachability,wyy-code/TA-OPD) — same acronym, different repo, different mechanism. REOPD (2608.11698) vs. the listed ReOPD (2607.04763) and REOPOLD (2603.11137). Resolve all by arXiv ID.
🎭 OPD with Black-Box / Outcome-Based Teachers
When the teacher is API-only (no logits), OPD uses scalar rewards, verbal scores, preferences, or adversarial discriminators — all evaluated on student rollouts. Entries that turned out to use static teacher data only (Lion, SuperCorrect, DAIL, SODA) are excluded from this list.
| Resource | 🌟 Stars | Date | Org | Paper Link | Title / Notes |
|---|---|---|---|---|---|
| ORPO-Distill | 2025.09 | Industrial | arXiv 2509.25100 | ORPO-Distill | |
LMOps /gad |
2025.11 | Microsoft Research | arXiv 2511.10643 · project | GAD — Black-Box OPD | |
| OVD | 2026.01 | HKU / Huawei | arXiv 2601.21968 | OVD (On-policy Verbal Distillation) — project page OVD.github.io 404s |
|
| SPoT | 2026.03 | Visual-AI | arXiv 2603.01683 | SPOT: Surgical Post-Training — black-box oracle edits student failures into proximal rollouts | |
| SODA | 2026.04 | Academic | arXiv 2604.03873 | SODA — Semi On-Policy Black-Box Distillation | |
| ROPD | 2026.05 | NUS / USTC / Tencent | arXiv 2605.07396 | ROPD — Rubric-based On-Policy Distillation; induces prompt-specific rubrics from teacher–student contrasts, then scores student rollouts by those rubrics (logit-free / black-box); up to 10× sample efficiency | |
| PRISM | 2026.04 | HKUST(GZ) / Tsinghua / NTU / RUC / USTC / UCAS | arXiv 2604.28123 | PRISM — Pre-alignment via Black-Box OPD for Multimodal RL; an adversarial OPD stage inserted between SFT and RLVR, scoring student rollouts with a Mixture-of-Experts discriminator (separate perception and reasoning experts, Bradley–Terry loss). +4.4 / +6.0 avg over SFT→RLVR at 4B / 8B | |
| OmniOPD | 2026.06 | Meta AI | arXiv 2606.01476 | OmniOPD — Logit-Free OPD via Speculative Verification; an entropy-driven scheduler picks uncertain chunks of the student’s rollout, a black-box teacher generates Monte-Carlo continuations, and they are scored by semantic similarity rather than logits. Beats white-box OPD with the same teacher family (69.08 vs 64.16) | |
| ExpRL | 2026.06 | Stanford / CMU | arXiv 2606.17024 | ExpRL — Exploratory RL for LLM Mid-Training; an LLM judge scores the student’s own rollouts and prefixes against a hidden reference solution under a fixed rubric, giving dense process-level reward before sparse RL ⚠️ see strictness note | |
| SOPD | 2026.08 | Nanjing Univ. (SKL) / XingYun Lab / UCAS / Fudan | arXiv 2608.16333 | Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning — at each step the teacher substitutes its own step for the student's, and the student is trained on the resulting relabelled trace; a step-granularity knob interpolates between pure OPD and pure SFT. Filed Black-Box because the teacher exposes no distribution: "the teacher supplies generated text only; SOPD requires neither teacher logits nor gradients". ⚠️ DAgger-style relabelling rather than scoring of the student's own continuation — see strictness notes. AIME24/25, HMMT25-Feb/Nov, ALFWorld |
📋 Click to view technical details
| Method | Feedback Signal | Data | Granularity | Domain | Notes |
|---|---|---|---|---|---|
| ORPO-Distill | Student-Generated Outputs (SGO) + ORPO contrastive | Mixed (student-generated negatives, teacher positives) | Sequence | Cross-architecture | "Mixed-policy strategy utilizing student-generated outputs"; NeurIPS 2025 WS. |
| GAD (Generative Adversarial Distillation) | Discriminator (on-policy reward model) | Student | Sequence | General | A trained discriminator distinguishes student outputs from teacher (e.g. GPT-5) responses; minimax game makes the discriminator co-evolve into an on-policy reward model. Qwen2.5-14B student becomes comparable to GPT-5-Chat on LMSYS. |
| OVD | Verbal scores (0–9) on student trajectories | Student | Sequence | General | Replaces token-level logit matching with verbal scoring; +25.7% over baselines. |
| SPOT | Black-box Oracle step edits + BCE reward objective | Student rollouts, Oracle-rectified | Step / sequence | Math reasoning | Minimal edits keep samples proximal to the student distribution, targeting reasoning gains with knowledge retention. |
| SODA | DPO: teacher responses as preferred vs. base student (q₀) zero-shot responses as rejected | Mixed | Sequence | Cross-architecture | "Semi on-policy" paradigm: captures student-specific inferior behaviors from a one-time static snapshot of q₀, eliminating the need for dynamic rollouts or adversarial training. 10× faster and 27% less peak GPU memory than GAD. Outperforms GAD on 15/16 benchmarks |
| ROPD | Prompt-specific rubric scores (Rubricator + Verifier) | Student | Sequence (rubric-weighted reward) | General | Black-box-compatible alternative to logit OPD: a Rubricator contrasts teacher vs. student responses to induce prompt-specific rubrics, a Verifier scores each student rollout into a weighted pass-rate reward. Needs only teacher-generated responses, no logits. |
📝 Strictness notes
-
ExpRL — ⚠️ The paper presents itself as RL mid-training, not distillation, and explicitly benchmarks against (and beats) a true OPSD baseline. It qualifies here only under the black-box reading: an LLM judge scores the student's own rollouts against a hidden reference under a fixed rubric. No teacher logits and no KL term anywhere.
-
OmniOPD — the teacher may be fully API-only; supervision is a continuous semantic-similarity score over Monte-Carlo teacher continuations at student-chosen chunks, not a distribution. Notable for outperforming white-box OPD with the same teacher family.
-
Excluded from this section after a full read: ZPPO (NVIDIA) — its own subtitle, "Teacher in Prompts, Not Gradients", states the disqualifier exactly: the teacher rewrites the prompt, the student resamples, and training is plain GRPO on binary reward. The teacher never scores a student token. It is the cleanest illustration of where C2 draws the line.
-
SOPD (2608.16333) — ⚠️ partially satisfies C2 in this section's own sense: the teacher exposes no distribution at all ("the teacher supplies generated text only; SOPD requires neither teacher logits nor gradients"), and at each step it substitutes its own step for the student's rather than scoring the student's continuation. That is DAgger-style relabelling; listed here because the relabelled states are still the student's own, and the step-granularity knob explicitly interpolates toward pure OPD.
♻️ Self-Distillation with Privileged Context — OPSD
Same model = teacher = student, but the teacher is conditioned on something the student doesn't see (verified trace, ground-truth answer, "be concise" prefix, longer context, document, …). The gap exists because of the conditioning, not weights.
Several entries previously listed here turned out on verification to use static teacher data or a fixed self-rewritten dataset rather than student rollouts; those have been excluded. SPIN was reclassified to Iterative Self-Bootstrapping.
| Resource | 🌟 Stars | Date | Org | Paper Link | Title / Notes |
|---|---|---|---|---|---|
| OPSD | 2026.01 | UCLA / Meta FAIR | arXiv 2601.18734 · blog | OPSD — Self-Distilled Reasoner | |
| Self-Distillation | 2026.01 | MIT / ETH | arXiv 2601.19897 | SDFT-Continual | |
| mtp-lm | 2026.02 | UMD / LLNL | arXiv 2602.06019 | MTP Self-Distill | |
LMOps /opcd |
2026.02 | Microsoft Research | arXiv 2602.12275 | OPCD — On-Policy Context Distillation | |
| GATES | 2026.02 | UMD | arXiv 2602.20574 | GATES (Self-Distillation under Privileged Context) | |
| EMPO² | 2026.02 | Microsoft Research | arXiv 2602.23008 · code · blog | EMPO² — memory-tip-conditioned online self-distillation for exploratory LLM agents (ICLR 2026; cross-listed into Agent) | |
| CRISP_Reasoning_Compression | 2026.03 | arXiv 2603.05433 | OPSDC / CRISP | ||
LMOps /oel |
2026.03 | Microsoft Research | arXiv 2603.16856 | OEL — Online Experiential Learning | |
| self-distillation-analysis | 2026.03 | MSR / KAIST / SNU | arXiv 2603.24472 | Why Does Self-Distillation (Sometimes) Degrade Reasoning? — diagnostic study of OPSD failure modes | |
| ml-ssd | 2026.04 | Apple MLR | arXiv 2604.01193 | Apple — Embarrassingly Simple Self-Distillation | |
| Skill-SD | 2026.04 | UCAS / CUHK / USTC / vivo AI Lab | arXiv 2604.10674 | Skill-SD — skill-conditioned OPSD for multi-turn LLM agents | |
| SD-Zero | 2026.04 | Princeton / Toronto / CMU | arXiv 2604.12002 | SD-Zero — Self-Revision turns binary rewards into dense supervision | |
| π-Play | 2026.04 | CASIA / UCAS / Meituan | arXiv 2604.14054 | π-Play — multi-agent self-play turns the question-construction path into privileged context for OPSD on search agents | |
| OPSDL | 2026.04 | Baidu | arXiv 2604.17535 | OPSDL (Long-Context Self-Distillation) | |
| MSD | 2026.05 | Tongji / Shanghai AI Lab | arXiv 2605.02971 | MSD — multilingual safety OPSD; teacher conditioned on English query translation + CoT instruction; DPSW weights safety-critical tokens | |
| COPSD | 2026.05 | LMU Munich / MCML | arXiv 2605.09548 | COPSD — crosslingual OPSD; teacher sees English problem translation + reference solution, student rolls out in low-resource language (17 African languages) | |
| SGSD | 2026.05 | THU | arXiv 2605.28791 | SGSD — Skill-Conditional Gated SD | |
| CODE | 2026.05 | USTC | arXiv 2605.28303 | CODE — OPSD on Knowledge Editing + Casual Editing | |
| SSOPD | 2026.05 | THU / Beihang | arXiv 2605.17497 | SSOPD — Self-Supervised OPSD; privileged context is the model's own shortest correct completion within a GRPO group (no external traces), distilled into prefixes of the longest wrong completion | |
| RLCSD | 2026.06 | THU (BPM) / Alibaba Tongyi | arXiv 2606.11709 | RLCSD — Contrastive OPSD; cancels privilege-induced style drift by contrasting the teacher–student gap under a correct hint vs. a wrong hint; verl-based | |
| d-OPSD | 2026.06 | THU / TUM / NTU / UT Austin | arXiv 2606.18195 | d-OPSD — first OPSD for diffusion LLMs; self-generated answers as suffix conditioning ("self future-experience"); step-level (not token-level) divergence aligned to the denoising process | |
| CaOPD | 2026.04 | Salesforce AI Research | arXiv 2604.16830 | CaOPD — The Illusion of Certainty; proves privileged conditioning makes the teacher's confidence a non-identifiable and upward-biased target, so OPD reliably buys accuracy at the cost of severe overconfidence; fixes it by rewriting confidence targets to the free empirical rollout success rate | |
| Vision-OPD | 2026.05 | ISCAS / UCAS / Xiaohongshu | arXiv 2605.18740 | Vision-OPD — regional-to-global OPSD for MLLM fine detail; teacher sees a 2×-upscaled evidence crop, student sees the full image. 9B beats Gemini-3.1-Pro on the fine-grained suite using 6.2K fully synthetic triplets, no GT labels or verifier | |
| D-OPSD | 2026.05 | HKUST / Z-Image Team Alibaba / UCSD / CUHK | arXiv 2605.05204 | D-OPSD — OPSD for continuously tuning step-distilled diffusion models; exploits that LLM/VLM-encoder T2I models inherit in-context ability, so feeding the encoder the target image is free privileged context | |
| PW-OPSD | 2026.05 | SaFo Lab / UW–Madison | arXiv 2605.21606 | PW-OPSD — When Are Teacher Tokens Reliable?; a branch-viability diagnostic shows position separates genuinely-uncertain from merely-diverse teacher tokens (AUROC 0.83) where every local uncertainty measure fails (≤0.57) | |
| OPSD Predictive Law | 2026.05 | Tufa Labs, Zürich | arXiv 2605.30070 | A Predictive Law for OPSD From World Feedback — the pre-training student–self-teacher accuracy gap linearly predicts final OPSD gain (R² 0.949 / 0.996), so privileged-context designs can be screened before training (ICML RLxF 2026) | |
| SDSD Diversity | 2026.06 | Mila / Univ. de Montréal / FAIR at Meta | arXiv 2606.26091 | OPSD with Sampled Demonstrations Reduces Output Diversity — proves the SDSD optimum tilts the base policy by pointwise conditional mutual information rather than reward, so it amplifies dominant modes; pass@1 up, pass@k flat | |
| PHF | 2026.06 | HKUST(GZ) / NUAA / NUDT | arXiv 2606.29340 | PHF — Privileged Hidden Flow; adds a residual-stream transition-geometry channel (direction + Gram-matrix CKA) on top of the standard OPSD output loss, with proven invariance to per-trajectory offsets and rescaling | |
| Visual-OPSD | 2026.06 | XJTU (MOE KLINNS) / SYSU | arXiv 2606.18974 | Visual-OPSD — shows a unified multimodal model's rendered "visual thoughts" matter as a generation pathway, not as pixels, then distills that pathway into a text-only student: +3.40 pp over its own teacher at 14.3× speedup | |
| Denser ≠ Better | 2026.07 | HKISI CAS / CASIA / UCAS / NJUST | arXiv 2607.01763 | Denser ≠ Better: Limits of OPSD for Continual Post-Training — SDPO specialises harder than GRPO but forgets far more (ToolUse −80% vs GRPO +18%); dense self-distillation is a specialisation accelerator, not a continual-learning stabiliser | |
| Purified OPSD | 2026.07 | ZJU / Tongyi Lab Alibaba / HUST / Jilin | arXiv 2607.02234 | Purified OPSD — Without Losing How to Think; decomposes the teacher update via a reference-only teacher and finds the reference-memorisation component dominates while the useful component is actively opposed (cos ≈ −0.95); re |