← Open Source
thinkwee

AwesomeOPD

Awesome List for On-Policy Distillation

ListsTool collectionsPaper collections
Open on GitHub
Momentum
+0stars in 24 hours0.0%
873
Stars
22
Forks
+0
This week
10
Contributors
Created 2026-04-27 · Updated 2026-10-01 · #12234 today
Top developers
README

banner Logo

Surveys White-Box Black-Box

OPSD Iterative OPD-RL

Reasoning Multimodal Agent

SpecDec Frameworks Industrial

When LLMs Distill On-Policy

AwesomeOPD is an awesome list summarising open-source repositories and papers for training LLMs (and VLMs / agents / draft models) with On-Policy Distillation (OPD) and On-Policy Self-Distillation (OPSD):

  • 🎯 OPD = C1 + C2. C1: student samples its own trajectories y ~ π_student(·|x) during training. C2: teacher provides per-token / sequence supervision on those student samples. Methods that only partially satisfy are flagged in 📝 Strictness notes per section.
  • 🪞 OPSD = special case where teacher is the same model, conditioned on privileged context (verified trace / answer / "be concise" prefix / longer context) or an earlier checkpoint.
  • 🚀 Each entry is annotated along four design axes — teacher source (external · same model with privileged context · earlier checkpoint · multi-teacher · discriminator), supervision signal (logits / top-k / sequence reward / verbal score / discriminator / verifier / feature), rollout consumption (all / selected / truncated / replaced / as PG samples), and pipeline slot (cold-start / mid / RL-replacement / inside-RL / inter-stage / compression / continual-anchor).
  • ⚠️ Built by reading paper PDFs, project pages, and source code with LLM coding agents; manually reviewed but errors possible. PRs welcome.
  • 📌 If you find this repository helpful for your research, please cite it via the "Cite this repository" button in the right sidebar of the GitHub page.
  • 📅 Last updated: 2026-08-26

Taxonomy:

  • 📚 Surveys, Foundations & Position Papers — meta-references and seed papers (GKD, MiniLLM, Thinking Machines blog, Tencent / THUNLP surveys)
  • 🔬 White-Box — logit-based OPD on student rollouts with an external teacher
  • 🎭 Black-Box — discriminator / verbal / preference, no teacher logits
  • ♻️ OPSD — privileged-context self-distillation (same model, different conditioning)
  • 🔁 Iterative Self-Bootstrapping — same model as previous-checkpoint teacher
  • 🤝 OPD-RL Hybrids — inside-RL OPD: KL-as-reward, RL+OPD fusion
  • 🧠 Reasoning / 🖼️ Multimodal / 🤖 Agent & Embodied — by application; cuts across all teacher-source categories
  • ⚡ Speculative-Decoding Distillation — drafter distillation; "student" is a draft model
  • 🛠️ Frameworks & Toolkits — what to actually run
  • 🏭 Industrial / Production Reports — what the labs ship

Shorthand: FKL = forward KL · RKL = reverse KL · JSD = Jensen–Shannon · Skew-KL / AKL = skewed / adaptive KL · 📄 paper-only = no public code yet.

Updates

📢 click to expand

  • 2026-08-26 — monthly sweep of 2026-07-20 → 2026-08-26: 107 new entries across every section, each read and then independently re-verified. Two passes: seven per-section sweeps proposed candidates against C1+C2, then ten adversarial verifiers re-checked every accepted entry against the arXiv abstract or HTML full text, attacking rather than confirming it. No entry survived on a proposer's word alone.
    • Headline numbers: OPSD +33 (the largest single-month addition the list has seen), Multimodal +27, White-Box +21, Agent +14, OPD-RL +10, Industrial +6, Surveys +6, Black-Box / SpecDec / Frameworks +1 each. Reasoning and Iterative Self-Bootstrapping had no qualifying new work. Zero duplicates against the 275 arXiv IDs already listed.
    • Rejected on the bar (5, after full-text reads): DreamMimic (2608.22278) — "imitation-dominant", teacher-driven rollout fraction, the FA-OPD failure pattern; SA-OPSD (2607.17136) — BC onto a verifier-corrected action, and dated one day outside the window; DiffusionOPSD (2608.24646) — frozen behaviour policy, reward-gradient targets, no privileged context. REOPD (2608.11698) and CROP (2608.13387) were both moved out of OPD-RL into White-Box: a PPO-shaped surrogate is not an RL loop when nothing but teacher log-probabilities enters the objective — the same reading under which G-OPD already sits in White-Box.
    • Metadata corrections found by verification: 41 Org cells the sweep had marked "unstated" were in fact stated in the arXiv HTML full text (among them Apple for Rubrics as Privileged Information, NIO for LS-MOPD, ByteDance for OPLD; StreamOPD's "unix-ai-lab" turned out to be a GitHub Pages slug, not an institution). 7 titles were wrong — five appended a parenthetical acronym the paper never uses, and two (FP-OPD, SPIRAL) were outright mis-stated. 3 entries marked 📄 paper-only have live repos (LOPD, SSPO, Veritas++), and Motif 3's weights turned out to be openly released. Only one "unstated" Org survived scrutiny (PCD, 2607.28336).
    • New acronym collisions (6): SMOPD ×2, RP-OPSD ×2, TA-OPD ×2, SPOT ×2, REOPD vs. ReOPD/REOPOLD, OPD-V vs. OPSD-V — all documented in the relevant sections' strictness notes.
    • The month's through-line: all six new Surveys entries are diagnostic or negative, and they converge independently on the same conclusion — the privileged information is doing less work than assumed. OP²SD swaps the reference for an unrelated problem's solution and stays competitive; Privileged Likelihood Is Not Automatically Value finds hindsight scoring near-random at AUC 0.505 with an outcome-only baseline beating every token-score variant; Test-Time Scaling argues OPD narrows the sampling budget rather than raising the ceiling and names it "illusory distillation". Read alongside the existing Rethinking OPSD for Thinking Models (2607.05184).
  • 2026-07-23 (c) — independent re-audit of all 191 entries added in (a) and (b). Every entry was re-read against its PDF by a second reviewer with no access to the first pass's reasoning; 189/191 verdicts were reproduced (98.9%).
    • Removed as out of scope (2): CollectionLoRA (2605.25378) — single-step LoRA image editing, no sequence model; FA-OPD (2605.27095) — MLP policy on Gym/D4RL continuous control, no language or generative-sequence component. The flow-matching entries that are retained (DiffusionOPD, D-OPSD, OPSD-V, dOPSD, AnyFlow, WMSD, π-Flow, Qwen-Image-2.0-RL) generate a supervised trajectory; these two do not.
    • Corrected: a transcription error in Multi-Rollout OPD (AIME25 mean@8 41.21 → 25.41); a misattributed ablation in OPSD Compresses RLVR (−26.8% → −29.22% at −1.03 pp); conflated settings in ADWIN; a worst-case-only figure in GeoSD; teacher count in MAD-OPD; a loose "doubles" in Tool-Call Boundary Drift.
    • Metadata: an unverifiable venue tag on Revisiting OPD; four wrong or over-specific Org cells (HPD listed a GitHub username; Draft-OPD listed an affiliation absent from the paper; CoPD "JD Explore" → JD.COM; EffOPD "Tencent Hunyuan" → Tencent); ten Date cells normalised to the arXiv-ID month; ShortOPD refiled from White-Box to OPSD (its teacher is its own pre-compression checkpoint, not an external model); two new acronym collisions documented (PBSD ×2, COPD/CoPD); all badge counts recomputed.
    • Checked and confirmed correct: the eight cross-listed arXiv IDs are the intended cross-section duplicates; HJSang/OPSD_OnPolicyDistillation genuinely hosts three separate papers (PACED, TIP, Sparse-to-Dense); no dead repo links.
  • 2026-07-23 (b) — backfill sweep of 2026-04-25 → 2026-06-20: 152 papers read in full, 120 added, 32 rejected. This period predates the previous update and had never been covered, so the list was missing roughly half the field's output during its busiest months. Highlights: Decoupling KL and Trajectories (the cleanest formal taxonomy of SFT/DAgger/offline-RL/OPD), Apple's Unmasking OPD, three independent parameter-geometry studies, the prefix-truncation family (Prune-OPD, ESR, ADWIN, Truncated OPD), cross-tokenizer OPD (SimCT, Tokenizer Barrier), Meta's logit-free OmniOPD, and two production reports with genuine MOPD (Kwai Keye-VL-2.0, Nemotron 3 Ultra). New acronym-collision and contested-finding notes added to the relevant strictness sections. Backfill entries carry their mechanism summary inline in the table rather than a separate technical-details row.
  • 2026-07-23 (a) — large sweep (~60 entries), every paper PDF read and every repo link checked.
    • Surveys: Formula-Driven Survey, NAIL, When Does Online IL Help, Demystifying OPD
    • White-Box: IW-OPD, DEAR, ReNIO, PG-OPD, SEAD, DOPD, MOPD (Xiaomi), Blockwise Drift Gating, Direct-OPD, OPD², TOP-D, COPD, ShortOPD, TOPD
    • OPSD: CaOPD, Vision-OPD, D-OPSD, PW-OPSD, OPSD Predictive Law, SDSD Diversity, PHF, Visual-OPSD, Denser ≠ Better, Purified OPSD, DemoPSD, Rethinking OPSD, GeoSD, AD-OPSD, dOPSD, PromptSD; BIRD → Iterative Self-Bootstrapping
    • OPD-RL: OPID, ATOD, PASS, CRAFT, UCOB, DRIFT, RG-OPD, CPO, Distilled RL, CriPO, H²SD, CADENCE
    • Multimodal: DiffusionOPD, V-Zero, H-OPD, OPSD-V, Med-OPD, OPD-IAD
    • Agent: KbSD, Two-Phase Distillation, SEED, ReOPD, UI-MOPD, TurnOPD, FA-SD Failure, Tool-Call Boundary Drift, ROAD-VLA
    • SpecDec: TIGER, AdaFlash · Frameworks: Spider, AsyncOPD, EasyOPD
    • Industrial: Solar Open 2, KAT-Coder-V2.5, OvisOCR2, Mach-Mind-4-Flash, Agents-A1, Qwen-Image-2.0-RL, Audex
    • Maintenance: TCOD upgraded to its released repo (COLM 2026); Draft-OPD relinked to the canonical Simplified-Reasoning org; venue tags added to HPD (ICML 2026) and Revisiting OPD (COLM 2026); all badge counts recomputed (several were stale)
    • Verified and excluded: TREK, REGEN, A Few Teacher Steps, Prefix-GRPO, HyperDFlash, SOPD-SocialNav, DataFlex-RL — each fails C1 or C2 on a full read; rationale in the relevant section's strictness notes
  • 2026-06-23 — add FiRe-OPD (White-Box; hard trajectory filtering + soft token reweighting)
  • 2026-06-19 — add d-OPSD, RLCSD, SSOPD (OPSD); TGPO, OPD+ (OPD-RL); Flow-OPD, Decomposed-OPD/VGS (Multimodal); ROPD (Black-Box); BRTS, OPRD (White-Box); Draft-OPD (SpecDec); SDAR (Agent); The Many Faces of OPD (Surveys)
  • 2026-06-13 — add TRD (White-Box; trajectory-level refinement) + Li Jiang's OPD reflection blog
  • 2026-06-08 — add EMPO² (OPSD; cross-listed into Agent)
  • 2026-06-07 — add SPOT
  • 2026-05-18 — add COPSD, MSD
  • 2026-05-15 — add TCOD, Healthcare AI GYM, HyperEyes (and cross-list Skill-SD into Agent)
  • 2026-05-14 — add CORD
  • 2026-05-13 — add Uni-OPD; add HY-MT, Baichuan-M3, KAT-Coder-V2, HY-Embodied, Qwen3.5-Omni
  • 2026-05-11 — add Skill-SD
  • 2026-04-30 — add π-Play
  • 2026-04-29 — add SD-Zero, Why Does Self-Distillation (Sometimes) Degrade Reasoning?
  • 2026-04-28 — initial release; add NPO, VLA-OPD, KDFlow, HPD, DeepSeek-V4

📚 Surveys, Foundations & Position Papers

Resource 🌟 Stars Date Org Paper / Link Title / Notes
GKD Paper 2023.06 Google DeepMind (Agarwal et al.) arXiv 2306.13649 — implemented in TRL GKDTrainer GKD: On-Policy Distillation of Language Models — Learning from Self-Generated Mistakes (Seminal · ICLR 2024)
Blog Blog 2025.10 Thinking Machines Lab (Kevin Lu et al.) Blog · tinker-cookbook Thinking Machines Lab — On-Policy Distillation (blog)
tinker-cookbook Stars 2025.10 Thinking Machines Lab — Reference impl. of the OPD recipe on the Tinker SDK
revisiting_opd Stars 2026.03 CASIA (Fu et al.) arXiv 2603.25562 Revisiting OPD: Failure Modes & Simple Fixes
Tencent OPD Survey Paper 2026.04 Tencent (Mingyang Song & Mao Zheng) arXiv 2604.00626 A Survey of On-Policy Distillation for LLMs
OPD Stars 2026.04 Tsinghua THUNLP arXiv 2604.13016 Rethinking On-Policy Distillation: Phenomenology, Mechanism & Recipe
Lightning OPD Paper 2026.04 Wu, Han, Cai arXiv 2604.13010 Lightning OPD: Efficient Post-Training with Offline OPD
OPSD Survey Paper 2026.05 Academic arXiv 2605.18141 A Brief Overview: On-Policy Self-Distillation in LLMs
Many Faces of OPD Paper 2026.05 UIUC (Ge Liu's ULab) arXiv 2605.11182 The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes — diagnoses when OPD/OPSD succeed or fail (distribution mismatch, optimization instability, PI-free policy limits) & proposes fixes; companion to Revisiting OPD & THUNLP Rethinking
Blog Blog 2026.06 Li Jiang Blog · arXiv 2606.08432 On-Policy Distillation: Promise, Pitfalls, and Prospects — reflection on OPD's promise, three failure mechanisms (local teacher noise, coverage decay, myopic gradients) & prospects; companion to TRD
Formula-Driven Survey Paper 2026.06 Tsinghua (Bowen Zhang) arXiv 2606.22793 A Formula-Driven Survey and Research Agenda for OPD — organises OPD as a feedback-to-update path rather than by KL direction; seven formula-derived variables + an explicit evidence-tier table (E0 papers → E3 own hypotheses)
NAIL Stars 2026.06 Columbia (Sriraman, Liu, Hsu, Block) arXiv 2606.30923 Behavior Cloning is Not All You Need: The Optimality of OPD for Noisy Expert Feedback — proves offline BC needs samples exponential in horizon under a noisy expert (and that this is necessary for any offline IL), while OPD is polynomial; proposes NAIL
Online IL Realizability Paper 2026.06 Tsinghua / CMU / Berkeley / Harvard arXiv 2606.30445 When Does Online Imitation Learning Help in LLM Post-Training? — challenges the error-accumulation story: under realizability OPD buys nothing over SFT; non-realizability is the real source of online gains
Demystifying OPD Paper 2026.07 CUHK / Tencent AI Lab arXiv 2607.13399 Demystifying OPD: Roles, Pathologies, and Regulations — OPD as exploration catalyst, not ceiling-raiser; diagnoses Student–Teacher Mismatch (the strongest teacher can be the worst) + Length Exploitation, and fixes both with clipping / log-compression
Unmasking OPD Paper 2026.05 Apple arXiv 2605.10889 Unmasking OPD: Where It Helps, Where It Hurts, and Why — training-free diagnostic that derives an ideal per-token gradient from empirical success probabilities and scores GKD / MiniLLM / Dr.GRPO objectives by cosine alignment to it. Finds guidance is far better aligned on incorrect rollouts than correct ones
Decoupling KL & Trajectories Paper 2026.05 EIT Ningbo / HK PolyU / HKUST / SJTU arXiv 2605.16826 A Unified Perspective for SFT, DAgger, Offline RL, and OPD — decomposes distillation along prefix source × KL direction, showing the four cells are exactly SFT / on-policy SFT / offline-RL distillation / OPD, and proves student-prefix + reverse KL equals a dense-reward REINFORCE objective
States, Not Tokens Paper 2026.05 Independent (Dong Nie) arXiv 2605.22731 Post-Training is About States, Not Tokens — recasts SFT/RL/OPD along state source × signal source; shows continuation-OPD from a deliberately degraded teacher still lifts the student past that teacher on all three axes
Geometry of OPD Paper 2026.06 HKUST / Brown / ZJU / HK PolyU / USTC / BUPT arXiv 2606.07082 On the Geometry of On-Policy Distillation — locates OPD between SFT and RLVR in parameter space (update sparsity 51.6% vs 8.1% / 77.2%) and identifies early subspace locking into a persistent low-rank update channel
Dense Supervision, Sparse Updates Stars 2026.06 Nanjing Univ. / Amap Alibaba arXiv 2606.13657 On the Sparsity and Geometry of OPD — OPD updates are tiny (0.036–0.142% relative Frobenius norm) and 67–90% coordinate-sparse, concentrated in FFN; training only the discovered sparse mask nearly recovers full OPD
Rock Tokens Stars 2026.05 UMBC / Case Western / ASU / VU Amsterdam arXiv 2605.09253 Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in OPD — tokens keeping high KL after saturation are mostly structural scaffolding whose gradients Adam neutralizes, and are causally near-irrelevant to accuracy
OPSD Compresses RLVR Paper 2026.05 Yonsei arXiv 2605.06188 OPSD Compresses What RLVR Teaches — separating correct-only from incorrect-only rollouts shows OPSD mainly compresses already-correct traces (−29.22% length at −1.03 pp accuracy, Qwen3-8B) rather than repairing failures; proposes it as a post-RL compaction stage
Extrapolation Cliff Paper 2026.05 NTU Singapore arXiv 2605.08737 The Extrapolation Cliff in OPD of Near-Deterministic Structured Outputs — derives a closed-form clip-safety threshold λ* beyond which reward-extrapolated OPD collapses output validity; the quantitative bound on G-OPD/ExOPD-style λ>1
Outcome-Confounded Supervision Paper 2026.07 CUHK Shenzhen (Guoqing Ma) arXiv 2607.23731 Outcome-Confounded Local Supervision in OPD — the standard local read of teacher/student token-level agreement vs. disagreement is confounded by the completed trajectory's outcome: "agreement-on-failure" tokens carry ~68% of token mass in math reasoning. Three candidate fixes were tested and none consistently resolves it; the paper is explicit that "Our contribution is therefore diagnostic rather than a new training method"
OP²SD Paper 2026.08 MBZUAI / Nagoya / RIKEN AIP arXiv 2608.09228 Privileged Solutions or Context-Induced Teacher Behavior? Dissecting OPSD — swaps the paired reference solution for one from an unrelated problem while holding student rollout, teacher and objective fixed; OP²SD stays competitive with standard OPSD, implying the gains come substantially from context-induced teacher behaviour rather than the privileged content itself. The sharpest ablation yet against the "learning from privileged information" story
Privileged Likelihood ≠ Value Paper 2026.08 Salesforce AI Research arXiv 2608.09263 Privileged Likelihood Is Not Automatically Value — three checks separating "does the score track better actions", "does feedback construction change what is compared", and "what does the loss actually reinforce". Under a common additive construction, hindsight-feedback scoring is near-random (AUC 0.505) and an outcome-only baseline beats every token-score variant (64.2% vs 24.2–33.9%, 20B model, AIME 2025)
Test-Time Scaling Lens Paper 2026.08 HKBU / SJTU / Shanghai Innovation Institute / Ant Group arXiv 2608.11829 Towards Understanding OPD through the Lens of Test-Time Scaling — reads OPD with pass@K / avg@K rather than pass@1: OPD narrows the sampling budget needed to reach a given answer instead of raising the reasoning ceiling. The pass@K advantage returns to the pre-OPD base model as K grows, and more previously-solvable problems become unsolvable than the reverse. Coins "illusory distillation"
Every Coin Has Two Sides Paper 2026.08 USTC / Peking / iQuest Research / MBZUAI / Zhejiang arXiv 2608.16647 On the Dual Nature of Generalization in OPD — systematic study across distribution shift, cross-domain transfer and multi-teacher setups; finds OPD "transfers a teacher's reasoning behavior rather than its answers to particular problems". Same-origin teacher/student pairs generalise across language and domain; cross-origin pairs mostly fit the training distribution, and multi-teacher combination creates capability trade-offs routing cannot fully control
Rethinking Privileged Information Paper 2026.08 FirstPrinciples arXiv 2608.18271 Rethinking Privileged Information in OPSD — on Qwen3 (1.7B–8B) across science/math, the correct privileged reference gives no consistent benefit: students improve without correct references, unrelated-problem solutions sometimes outperform correct ones, and student predictions track the base model more closely than the reference. Companion negative result to OP²SD and to Rethinking OPSD for Thinking Models
Fin-SD ablations Blog 2026.09 CTGT (Yu, Gorlla) Blog Privileged Information in Self-Distillation: An Empirical Study — ablates failure-point hint self-distillation (Fin-SD) on finance reasoning with GPT-OSS-20B. Replacing the entire hint with a content-free placebo ("Your mistake is conceptual") matches the full hint (72.52% vs 72.19%, base 69.16%); a same-family hint writer ~6× the student's size adds nothing (both 72.19%); the hint variants match full-solution OPSD (72.19% vs 70.42%) with 9.5% fewer eval tokens. Zero-content version of the unrelated-solution control in OP²SD and Rethinking Privileged Information

📋 Click to view technical details

Resource Loss / Divergence Data Teacher Access Granularity Notes
GKD (Agarwal) Generalised JSD (FKL/RKL configurable) Mixed (λ interpolates teacher↔student) White-box Token The seminal paper that named OPD; introduced student-self-rollout supervision.
Thinking Machines blog Reverse KL (student‖teacher) Student rollouts White-box Token "Swap KL ref model for stronger teacher" recipe; one-line addition to RL trainer. Replicates Qwen3 result at ~1/10 RL cost.
Revisiting OPD Truncated reverse KL + top-p sampling + special-token masking Student White-box Token (filtered) Diagnoses 3 failure modes: imbalanced one-token signal, unreliable prefix guidance, tokenizer mismatch.
Tencent OPD Survey (survey) (survey) (survey) (survey) Catalogues 50+ methods; useful as a reference index.
THUNLP Rethinking OPD Reverse KL with progressive top-K alignment Student White-box Token Identifies two success conditions: compatible thinking patterns + genuinely new teacher capability. Recipe = off-policy cold-start + teacher-aligned prompt selection.
Lightning OPD Cached teacher log-probs over SFT rollouts (offline OPD) Student (cached) White-box Token Introduces "teacher consistency" — same teacher must be used for SFT and OPD or else gradient bias. Eliminates the live teacher server.
OPSD Survey (survey) (survey) (survey) (survey) Categorize eight designs; useful as a reference index.
Many Faces of OPD Reverse KL (FKL/RKL comparison) Student (self-sampled trajectories) White-box / self Token Diagnostic study (not a training method). Identifies three failure mechanisms — distribution mismatch, optimization instability, and a PI-free policy limit specific to OPSD — and proposes stop-gradient objectives + stabilized training as fixes. Covers both OPD and OPSD.
Formula-Driven Survey (survey) (survey) (survey) (survey) Refuses the "KL direction × teacher access" taxonomy; models OPD as feedback→update with seven variables: state distribution, feedback source, comparison support Ω_t (sampled-token / top-k / full-vocab), temporal credit A_t, gate-or-weight w_t, vocabulary routing, update route (direct-loss vs policy-gradient score-function). Carves out OPD-hybrids and OPD-adjacent (RLVR, teacher-forced KD, SFT) as out of scope. Proposes two unimplemented designs (GAE-OPD, CR-OPD), self-labelled as E3 hypotheses.
NAIL (Behavior Cloning Is Not All You Need) Forward-KL and reverse-KL variants, on a rollout distribution distinct from the scored policy Student White-box Token Theory + method. Noisy-expert model π*_η = (1−η)π* + η·ν. Offline BC needs n ≳ (1−η)^{−(H+2)}log|Π|/ε — exponential in horizon and necessary for any offline IL algorithm, unlike the horizon-free clean-expert result; the OPD variant is polynomial in H. Key design claim: the loss must distinguish the rollout distribution (greedy) from the scored policy (temp 1) — which standard OPD does not. NAIL ≈ BC/OPD at low noise, far better at high noise (GSM8K + modular addition).
When Does Online IL Help Reverse KL + entropy under the student (analysis only) Student White-box / self Sequence (contextual bandit, H=1) Position/theory paper. Under realizability (student class can represent the expert) SFT on expert samples fully matches expert performance and OPD adds neither accuracy nor speed — replicated on Countdown, GSM8K (Llama-3.2-3B) and DeepScaleR (R1-Distill-Qwen-1.5B). Under misspecification, discrepancy-based (TV/Hellinger) bounds go vacuous, and an information-theoretic lower bound with coverage coefficient C^e_∞ applies even at H=1. Notes strong-to-weak OPD lives in the non-realizable regime — which is why it works.
Demystifying OPD Reverse KL as policy gradient, per-token Δℓ_t = log π_T − log π_θ, PPO-clipped Student White-box Token Analysis + method. (1) OPD steers the student toward correct paths without expanding the capability ceiling (pass@1024 converges to the base model's); prompt diversity beats per-prompt rollout depth. (2) Student–Teacher Mismatch: Qwen3-4B-GRPO, the strongest teacher, is the worst teacher (student stuck ~2% AIME25); an Informativeness metric I = E[Δℓ̄|r=1] − E[Δℓ̄|r=0] predicts this in advance. (3) Length Exploitation: token-mean advantage lets the student wash out negative advantage by padding or truncate early on a favourable prefix. Fixes = Hard Clipping + Soft Log-Scale Compression, zero extra compute. Regulated OPD with a 1.7B/4B teacher beats prior OPD from a 30B teacher (AIME24 45.2 vs Uni-OPD 35.2 / G-OPD 37.3).
Fin-SD ablations (CTGT) Per-token reverse KL over a 100-token window from the located error (beat forward KL in this short-horizon setup) Student (on-policy prefix, truncated just before the first incorrect step) Self (student conditioned on a short hint at the error; hint written by a frozen student copy or a ~6× same-family model, or replaced by a fixed placebo) Token Empirical study. Hint content and hint-writer size both ablated with no accuracy effect at 20B; windowing 8192→100 tokens improves accuracy, completion and length; residual style-token drift (more "but / maybe / however") and length growth with more epochs. Stated scope: finance reasoning at ≥ GPT-OSS-20B capability, where the needed concept is already latent.

📝 Strictness notes (against the strict OPD definition C1: student samples its own trajectories during training + C2: teacher provides supervision on those samples)

  • Lightning OPD — ⚠️ partially satisfies C1: teacher log-probs are pre-computed once over SFT rollouts and reused during training; student doesn't actively sample during the OPD step. Authors call this "offline OPD" explicitly. Listed in OPD because the data is past-student-generated rollouts, not teacher-generated.

🔬 OPD with Larger External Teachers — White-Box

White-box methods use teacher logits / log-probabilities to supervise the student on student-generated rollouts. Each entry below has been verified to (a) train on student rollouts and (b) operate at the token level.

Methods that turned out to be RL-style on verification have been moved to OPD-RL Hybrids; off-policy / pure-loss-function / pretraining-side methods are excluded from this list.

Resource 🌟 Stars Date Org Paper Link Title / Notes
LMOps /minillm Stars 2023.06 Microsoft / Tsinghua arXiv 2306.08543 MiniLLM (ICLR 2024)
distillm Stars 2024.02 KAIST / Microsoft arXiv 2402.03898 DistiLLM (ICML 2024)
google-research /speculative_kd Stars 2024.10 UCSB / Google arXiv 2410.11325 Speculative KD (ICLR 2025)
distillm-2 Stars 2025.03 KAIST / Microsoft arXiv 2503.07067 DistiLLM-2 (ICML 2025 Oral)
DSKDv2 Stars 2025.04 BJTU arXiv 2504.11426 DSKDv2 — cross-tokenizer; supports on-policy mode
Constrained OPD Paper 2025.09 Huawei Noah's Ark arXiv 2509.22921 Constrained OPD (CMDP)
AdaSwitch Paper 2025.10 RUC / Baidu arXiv 2510.07842 AdaSwitch (on-/off-policy switching)
Veto Paper 2026.01 SNU arXiv 2601.07155 Veto (Stable OPD) — ACL 2026 Findings
G-OPD Stars 2026.02 RUC / Tencent arXiv 2602.12125 G-OPD
Fast OPD Paper 2026.02 Industrial arXiv 2602.15260 Fast OPD (prefix-truncated)
Entropy-Aware OPD Paper 2026.03 KAIST / IBM arXiv 2603.07079 Entropy-Aware OPD
REOPOLD Paper 2026.03 KAIST / Microsoft arXiv 2603.11137 REOPOLD (Relaxed OPD) — code soon
OPSD_OnPolicyDistillation Stars 2026.03 LinkedIn arXiv 2603.11178 PACED — frontier curriculum self-distill
TSD-KD Stars 2026.03 Korea Univ. arXiv 2603.13260 TSD-KD — token-selective dual KD (ICLR 2026)
SCOPE Stars 2026.04 USTC / Meituan / Fudan arXiv 2604.10688 SCOPE — signal-calibrated dual-path
OPSD_OnPolicyDistillation Stars 2026.04 Meta / LinkedIn arXiv 2604.14084 TIP — Token Importance, shares LinkedIn OPSD repo with PACED
Hybrid-Policy-Distillation Stars 2026.04 SJTU / Shanghai Innovation Institute / Tencent arXiv 2604.20244 HPD — Hybrid Policy Distillation; LlamaFactory + verl backends (ICML 2026). ⚠️ headline Tables 2–4 use a lightweight approximation of on-policy sampling that avoids full-sequence rollouts; true on-policy results appear only in §5.4/Table 5
BRTS Stars 2026.05 JHU (Patel group) arXiv 2605.09725 BRTS — Best-of-N Teacher Rollout Selection; augments student-context OPD with a curated teacher-context branch (correctness-first, then student-alignment) to cut single-rollout teacher variance
FiRe-OPD Stars 2026.06 THU / HKUST / Meituan (Li et al.) arXiv 2606.02684 FiRe-OPD — Filter, then Reweight; decouples optimization granularity — hard trajectory-level filtering (drop bottom-p% rollouts by teacher log-prob) + soft token-level reweighting (teacher-confidence × student-confusion), arguing soft weighting beats hard token selection (cf. TIP); verl-based, with a multi-teacher math+code variant
trd Stars 2026.06 McGill / Mila / UT Austin (Jiang et al.) arXiv 2606.08432 TRD — Trajectory-Refined Distillation; diagnoses prefix failure of dense per-token OPD, refines student rollouts at trajectory level before distilling; verl-based, also applies to OPSD
OPRD Stars 2026.06 ZJU / Ant Group arXiv 2606.06021 OPRD — On-Policy Representation Distillation; first OPD to supervise in hidden-state space (aligns teacher/student representations across layers on student rollouts, bypassing the LM head) rather than logits; built on the THUNLP OPD stack
IW-OPD Stars 2026.06 Xidian / Georgia Tech / Amazon AGI SF Lab arXiv 2606.22600 · project IW-OPD — On the Position Bias of OPD; supervising only the prefix 30% of tokens matches full-token OPD while suffix-30% barely learns; reweights the OPD advantage by a prefix-importance term derived from a trust-region argument
DEAR Paper 2026.06 Meituan LongCat / Nanjing Univ. / TJUNLP arXiv 2606.22830 DEAR — Finding the Evidence; argues entropy-selective OPD (cf. TIP) captures only decision tokens and structurally misses low-entropy, high-divergence evidence tokens where the student is confident yet wrong
ReNIO Stars 2026.06 ECNU / Shanghai Innovation Institute arXiv 2606.23104 ReNIO — Reweighting Negative Trajectory Importance; finds training on incorrect student outputs beats correct-only under both OPD and OPSD, then approximates correctness with a prefix-computable pivotal-token proxy
PG-OPD Paper 2026.06 TeleAI / SJTU arXiv 2606.21994 Prefix-Guided OPD — Mining Golden Trajectories from Rollouts; probes candidates at a short prefix by teacher–student top-k overlap and only continues the promising ones to full length (up to +4.80 avg, 2.46× wall-clock)
SEAD Paper 2026.06 Capital One arXiv 2606.28562 SEAD — Competence-Aware OPD; entropy zones the tokens (skip / RKL / FKL), cosine-anneals FKL→RKL, and adds the first prompt-level curriculum for OPD
DOPD Paper 2026.06 JD Explore / NUS / MMLab CUHK / PKU arXiv 2606.30626 DOPD — Dual On-policy Distillation; names "privilege illusion" (part of a privileged teacher's gap is information asymmetry, not transferable capability) and routes each token among four regimes by a privilege-advantage gap
MOPD Paper 2026.06 Xiaomi LLM-Core / PKU / HKU / RUC arXiv 2606.30406 MOPD — Multi-Teacher OPD for Capability Integration; the method paper behind MiMo-V2-Flash's MOPD stage. Shows same-origin teachers are essential — a stronger but distributionally distant teacher diverges outright
Blockwise Drift Gating Paper 2026.06 Independent (Zheng & Jiang) arXiv 2606.24084 Blockwise Policy-Drift Gating — student-only old↔current drift gate for OPD under rollout reuse (author-labelled preliminary study)
Direct-OPD Stars 2026.07 Tsinghua AIR SIA-Lab / ByteDance Seed / PKU arXiv 2607.05394 · project Direct-OPD — Weak-to-Strong Generalization; distills the teacher's implicit reward (post-RL minus pre-RL log-ratio) instead of its distribution, so a weaker post-RL teacher can improve a stronger student. Qwen3-1.7B AIME24 48.3→58.3 in ~4 h on 8×A100
OPD² Stars 2026.07 NAVER AI Lab arXiv 2607.15161 On-Policy Delta Distillation; replaces the teacher–student log-ratio reward with a delta signal (teacher minus its own base model), plus a sign-agreement gate. Qwen3-1.7B math avg 34.8→54.6
TOP-D Paper 2026.07 HKUST(GZ) / Microsoft arXiv 2607.04751 Trust Region Policy Distillation; interpolates teacher and student in probability space so the token reward is lower-bounded and gradient variance is provably bounded. AIME24 avg@32 50.42 vs 24.58 for standard OPD
COPD Paper 2026.07 SJTU / Alibaba Qwen arXiv 2607.19046 Contrastive OPD; the teacher re-scores the student's tokens twice — under a "light-thinking" and a "heavy-thinking" prefix — and the contrast becomes the advantage. +2.1 pp with 57% shorter responses; ships a COPSD self-distillation variant
TOPD Paper 2026.07 CASIA / UCAS arXiv 2607.16872 Trace-Based OPD for masked diffusion LMs; supervises only trace-aligned denoising decisions (commitments surviving into the final answer), avoiding the backward-reconstruction states random-mask supervision creates. MATH500 +5.7 with 4× fewer rollout rounds
AOPD Paper 2026.05 HUST / PKU / Meituan arXiv 2605.06387 Asymmetric OPD; splits the OPD gradient into positive (exploit) and non-positive (imitate) token regions, replacing the noisy policy gradient with truncated forward KL only on the latter. Preserves math accuracy after continual tool-use training (+0.61 vs −7.49 for OPD)
SimCT Stars 2026.05 Tencent Hunyuan / USTC / Shanghai Innovation Institute arXiv 2605.07711 SimCT — Recovering Lost Supervision for Cross-Tokenizer OPD; builds a common supervision space of shared-vocab tokens ∪ minimal aligned units (finest jointly-tokenizable spans) so the unchanged OPD loss applies across tokenizers
Prune-OPD Stars 2026.05 HKUST(GZ) / MBZUAI / UC Merced / SYSU arXiv 2605.07804 Prune-OPD — monitors per-position teacher/student top-k overlap, attenuates rewards on prefix drift and truncates once supervision is locally unexploitable; −37.6–68.0% training time. The mechanism KAT-Coder-V2.5 credits for its drift-aware truncation
ESR Paper 2026.05 UCLA / BIGAI arXiv 2605.27028 Less is More: Early Stopping Rollout for OPD; names Off-Policy Teacher Decay — conditioning the teacher on the student’s drifting prefix collapses its late-token distribution toward student-level autocomplete. Beats full-rollout OPD in every cell at up to 24× lower cost
ADWIN Paper 2026.05 PKU / Tencent arXiv 2605.28396 ADWIN — Adaptive Windows for Horizon-Aware OPD; truncates by a prefix-admissibility criterion based on gradient-cosine alignment between short-prefix and full-rollout updates, audited by delayed full-rollout probes. 59.3→60.9 avg at ~3.4× fewer FLOPs single-task (the 4.1× figure is from the separate strong-to-weak setting)
Truncated OPD Paper 2026.05 CASIA / Meituan arXiv 2605.31490 Are Full Rollouts Necessary for OPD? — proves sequence-level OPD accumulates noisy future signal at O(T³) MSE while token-level does not; truncating to 10% of the horizon matches full-rollout OPD while cutting training time 82%
Prefix Teach, Suffix Fade Paper 2026.05 ZJU / Meituan LongCat / Jilin arXiv 2605.13643 Local Teachability Collapse in Strong-to-Weak OPD; cuts dense supervision at a BIC-detected change-point in the teacher’s top-1/top-2 margin over the student’s reachable candidates. Avg 36.8→40.1
KAT Paper 2026.06 HKUST(GZ) / HKUST / HK PolyU / EIT arXiv 2606.09471 Escaping the KL Agreement Trap in OPD — low KL is not evidence of health: the student drifts into a corrupted prefix and the teacher locally agrees, silencing the corrective signal. Sliding-window detection cuts rollout length 59.7% at 2.4× speedup
TA-OPD Stars 2026.05 HK PolyU / InfiX.ai arXiv 2605.26844 Not All Disagreement Is Learnable — Token Teachability in OPD; supervises only a support-aligned "teachable" subset, matching or beating full-token OPD with 5–10% of tokens
TRB Paper 2026.05 T-Tech arXiv 2605.31159 Trust-Region Behavior Blending; leaves the OPD loss untouched and instead controls the rollout behaviour policy μ ∝ π_S^(1−β)·π_T^β under a trust-region constraint, annealing back to pure on-policy sampling
TrOPD Stars 2026.06 Samsung Research Beijing / Oxford / PKU arXiv 2606.01249 Trust Region On-Policy Distillation; partitions student tokens into a teacher-verifiable trust region (reverse KL) vs outliers (top-k forward KL), plus annealed off-policy guidance. AIME25 44.06 vs OPD 40.72
Near-Future Guidance Paper 2026.06 UMBC arXiv 2606.00305 Bridging Reasoning Trajectories in OPD via Near-Future Guidance; shows per-token KL is a weak proxy for real trajectory divergence (Pearson r=0.126) and adds an optimal-transport-aligned target spanning several future positions
PowerOPD Paper 2026.06 EIT Ningbo / HK PolyU / SJTU / Waterloo arXiv 2606.17199 Stabilizing OPD with Bounded Power Transformation; the vanilla reward log(π_T/π_θ) is unbounded and drives instability, so it is replaced by a sign-consistent Box-Cox transform bounded in [−1,1]. +6.37 Avg@8 at 59% less wall-clock
NPD Paper 2026.05 Huawei / Tianjin Univ. arXiv 2605.05940 Near-Policy Distillation; decouples generation from updates so teacher top-k logits can be computed by packed parallel prefill — 8.1× speedup, at the cost of deliberate policy lag
RWOPD Paper 2026.05 NUS arXiv 2605.13501 Reward-Weighted OPD with an Open Property-Equivalence Verifier; forward KL on rollouts passing a SymbiYosys+Z3 equivalence check, for NL→SystemVerilog Assertions. Beats its own 14B teacher and 671B general baselines
ProteinOPD Stars 2026.05 Tsinghua / IDEA / HKUST(GZ) / NTU arXiv 2605.10189 ProteinOPD — multi-teacher OPD for protein language models; distills a normalized product-of-experts consensus over preference-specific teachers via token-level generalized JSD. PPL −83.7%, thermostability +54.2%
MAD-OPD Stars 2026.05 HUST / Alibaba arXiv 2605.01347 Breaking the Ceiling in OPD via Multi-Agent Debate; two larger teachers (K=2) debate across rounds to form the token-level target, with task-adaptive divergence. A 4B student beats its own 14B teacher on LiveCodeBench-v6
W2S-OPD Stars 2026.07 UMD / Microsoft Research / MBZUAI arXiv 2607.26246 · project W2S-OPD — Weak-to-Strong On-Policy Distillation; synthesises a proxy teacher distribution softmax(z_base + α(z⁺ − z⁻)) by re-anchoring a weak contrast pair's logit delta onto the student's own base model, then distils it with per-token reverse KL. Three interchangeable contrast sources: post-RL vs pre-RL expert, larger vs smaller base model, correct vs wrong hints
Cross-Tokenizer OPD Paper 2026.07 UCAS / KwaiKAT Team / Zhejiang Univ. arXiv 2607.22334 Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization — maps the teacher's full next-token distribution into the student's vocabulary by marginalising over byte prefixes, so the unchanged β-divergence (FKL/JSD/RKL) applies across mismatched tokenizers with no hand-built alignment table. Third cross-tokenizer construction alongside SimCT and SimpleOPD
Relay-OPD Stars 2026.07 Zhejiang Univ. / Yuvion Team, Alibaba arXiv 2607.26057 Pass the Baton: Trajectory-Relayed On-Policy Distillation — attacks prefix failure via a teacher–student continuation asymmetry on failed prefixes (the teacher redirects, the student ploughs on) used as a label-free handoff trigger: the teacher briefly writes a "teacher leg", the student resumes, and the relayed trajectory is distilled under a limited relay budget concentrating intervention early. +5.73% over OPD, +1.49% over FastOPD. Same diagnosis as TRD, different remedy (in-trajectory relay vs. post-hoc refinement)
TU-OPD Paper 2026.07 Moore Threads AI arXiv 2607.27770 Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold — multi-teacher OPD that quality-gates each teacher per rollout: a weight α_i(r) selects which specialist's distribution supervises which student sample, plus a residual term preserving specialist behaviour instead of averaging it away. Verifier score only weights/selects teachers; it never enters the objective
Flux-OPD Paper 2026.07 PKU / Kling Team / Tsinghua / SJTU / Zhongguancun Academy arXiv 2607.28022 Flux-OPD — On-Policy Distillation with Evolving Contexts — for open-ended domains where rewards are unverifiable: a context expressing the desired preference is fed to the teacher and the difference between context-conditioned and context-free teacher distributions becomes the signal, with the context itself evolving across iterations. Sibling of OPD²'s delta idea, but contrasting the same teacher under two conditionings — two loaded models instead of three
Adaptive FastOPD Paper 2026.07 USTC / Shanghai AI Lab / SJTU arXiv 2607.29494 Adaptive FastOPD — Progress-Aware Rollout Horizon Expansion — makes the prefix-truncation family adaptive: rather than a fixed truncation length the rollout horizon expands as measured reasoning progress warrants, recovering full-rollout quality at truncated-rollout cost. Direct successor to Fast OPD; belongs with Prune-OPD / ESR / ADWIN / Truncated OPD
OPTD Paper 2026.08 HKUST / NPU / HK PolyU arXiv 2608.02942 OPTD — On-Policy Transition Distillation for Few-Step Diffusion LMs — samples partial states from the student's own denoising trajectory and matches transitions against a frozen teacher, with a set-bottleneck target and consistency-guided adaptive compression. ⚠️ Same dLLM caveat as TOPD: "trajectory" is a denoising path over masked positions, not left-to-right generation
SA-OPD Paper 2026.08 Zhejiang Univ. / ByteDance / Shanghai AI Lab arXiv 2608.03632 When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation — a teacher's confident tokens are not always input-grounded; some reflect priors rather than the actual problem. Filters those spurious tokens out of the OPD signal by an input-groundedness gap test (Δ_IG threshold) before applying reverse KL. Only the LLM track is in scope; the paper carries a second VLM track
SPOT (sparse probing) Paper 2026.08 CUHK-Shenzhen / ECNU / Meituan / Beihang / CityU HK arXiv 2608.04419 SPOT — Sparse Probing and Outcome Calibration — cuts teacher cost by sparsely probing only informative positions of the student rollout (chosen by teacher entropy / top-k candidate mass) and calibrating the resulting targets against the trajectory outcome, giving a closed-form KL-regularised target across acquisition / exploration / exploitation phases. ⚠️ Acronym collision with the Black-Box SPOT (2603.01683, "Surgical Post-Training") — unrelated; resolve by arXiv ID
OPD² (multilingual) Stars 2026.08 NAVER AI Lab arXiv 2608.05802 On-Policy Delta Distillation for Multilingual Math Reasoning — carries the OPD² delta signal log π_teacher − log π_base into cross-lingual math (EN/KO/JA), where the teacher's absolute distribution is contaminated by language priors the student already holds. ⚠️ Shares both the OPD² name and the same repo as 2607.15161 — a companion second paper, not a replacement; the existing note about needing the teacher's pre-post-training base checkpoint applies here too
Simple-OPD Paper 2026.08 Tsinghua / HKU / Tencent (LLM Dept.) arXiv 2608.06802 Simple-OPD — Demystifying Warm-up for On-Policy Distillation — asks what the SFT warm-up before OPD actually buys and finds a cheap LoRA warm-up suffices, after which plain reverse-KL OPD matches far more elaborate recipes. Speaks directly to the "off-policy cold start" finding in THUNLP's Rethinking OPD
WDL-OPD Paper 2026.08 Beihang / China Telecom eSurfing Cloud arXiv 2608.09447 WDL-OPD — Weak-Driven OPD via Mixture-Constrained Co-Training — only a designated anchor model rolls out ("Only the anchor generates trajectories: y ∼ π_A(·|x)"); the supervision target is a geometric mixture of the anchor's and an auxiliary model's token distributions, constrained toward a frozen teacher — co-training two students against one teacher rather than distilling each separately
TIDE Stars 2026.08 HKU / USTC / CUHK arXiv 2608.09836 Mismatch Matters: OPD Beyond Token Agreement — splits teacher–student mismatch into excess mass (student over-commits) and deficit mass (student never proposes what the teacher wants), penalising the first with a Hellinger-shaped term and repairing the second by injecting teacher top-K tokens, arguing plain token agreement hides both failure modes. verl-based
ReOrder-OPD Paper 2026.08 Hello Group / Institute of Automation CAS / UCAS / PKU arXiv 2608.10905 ReOrder-OPD — Reliability-Aware Prompt Ordering — leaves the OPD loss untouched and reorders the prompt curriculum by how reliable the teacher's supervision is on each prompt; a data-ordering counterpart to SEAD's competence-gated curriculum. The ROUGE proxy orders prompts only — the token-level target uses real teacher logits
REOPD Paper 2026.08 Shanghai AI Lab / PKU / SWJTU / USTC / UESTC / RUC arXiv 2608.11698 REOPD — Reliability-Adaptive Reward Extrapolation — generalises ExOPD/G-OPD reward extrapolation from a single global λ to a per-token compatibility weight gating a batch-adaptive extrapolation budget, cleanly separating base alignment from the extrapolation residual. Filed here rather than in OPD-RL: the advantage traces entirely through log π_θ / log π_T / log π_ref, and the paper states it "requires no verifier, reward model, value model, or extra rollout beyond standard OPD". ⚠️ Acronym near-collisions with ReOPD (2607.04763) and REOPOLD (2603.11137)
CROP Paper 2026.08 University of Hong Kong arXiv 2608.13387 CROP — Task Relevance via Counterfactuals for Selective OPD — scores each token's task relevance by counterfactual intervention (would the outcome change if this token changed) and distils only the tokens that matter. Adds to the token-selection debate alongside TIP / DEAR / TA-OPD / R2-OPD. Mechanism is a hard token mask over the unchanged clipped OPD surrogate — no reward or verifier signal, hence White-Box rather than OPD-RL
SimpleOPD Stars 2026.08 Shanghai AI Laboratory (SU-01 Team) arXiv 2608.14277 SimpleOPD — Tokenizer-Agnostic OPD for Long-Context Reasoning — matches token log-probabilities across differing vocabularies with no alignment table, distilling a long-context teacher into a short-context student. Third cross-tokenizer entry alongside SimCT and 2607.22334, each with a different alignment construction
DUET Paper 2026.07 WeChat, Tencent Inc. arXiv 2608.14644 DUET — Dual-Teacher OPD via Same-Weight Disagreement — pairs two teachers with identical weights, one shown the prohibition and one not, and treats their token-level disagreement as an isolated, noise-free supervision signal for runtime-injected policy compliance. 72.3–85.2% violation compliance at 88–93% retained utility. Same same-weight-different-context contrast as COPD and Flux-OPD, applied to safety
AED Stars 2026.08 Sun Yat-sen Univ. / Hong Kong Metropolitan Univ. arXiv 2608.14685 Rethinking Reverse KL as Adaptive Entropy Distillation — decomposes the reverse-KL OPD objective into a teacher-fitting term and a student-entropy term, then uses the teacher's own per-token entropy to calibrate imitation strength: high-entropy (uncertain-teacher) tokens get less forced imitation, low-entropy tokens more — with no explicit forward-KL branch. Full-vocab and teacher-top-k variants; MATH-500 Avg@8 14.63%→47.80% over vanilla RKL in one reported config
TA-OPD (tail-aware) Stars 2026.08 SUSTech (Statistics & Data Science) arXiv 2608.14728 Tail-Aware Top-k On-Policy Distillation — top-k teacher truncation silently discards the tail probability mass; TA-OPD keeps the aggregated tail as an explicit extra distillation target, in three variants (normalised top-k, TA-OPD, sample-corrected TA-OPD). verl/vLLM on DAPO-MATH-17K. ⚠️ Acronym collision with the listed TA-OPD (2605.26844, token Teachability, repo wyy-code/TA-OPD) — different paper, different repo, different mechanism
Open-MOPD Stars 2026.08 SIA-Lab (Tsinghua AIR & ByteDance Seed) / Tsinghua arXiv 2608.19098 Open-MOPD — Diagnosing and Fixing Capability Imbalance in Multi-Teacher OPD — the open reproduction of the Xiaomi MOPD recipe: diagnoses naive multi-teacher merging (math / code / instruction-following) as an optimization-budget imbalance rather than a capability conflict, and rebalances it — integration gap 3.50 → 0.31 points, recovering 83.4% of the achievable maximum on SmolLM3-3B. Ships 5 checkpoints + 3 domain RL teachers + data + full pipeline, making it the most reproducible multi-teacher OPD artifact in the list
R2-OPD Paper 2026.08 HKUST(GZ) / Tsinghua / Zhejiang / Univ. of Alberta arXiv 2608.19408 Beyond Imitation: Filtering OPD by Reasoning Progress — argues teacher likelihood and actual reasoning progress can point in opposite directions, and masks out tokens where the teacher-imitation ranking disagrees with an independently estimated progress reward, keeping only supervision that moves the solution forward. A non-entropy, non-divergence selection criterion for the token-selection debate

📋 Click to view technical details

Method Loss / Divergence Data Granularity Domain Notes
MiniLLM Reverse KL via policy gradient Student Sequence (PG) General The seminal "OPD" recipe by Yuxian Gu et al.; predates GKD by days. Mode-seeking.
DistiLLM Skewed-KL (mix of FKL/RKL) Mixed (adaptive off→on, with student samples) Token General Skew parameter α interpolates between FKL and RKL; importance-reweighted student samples.
Speculative KD (Xu) Interleaved propose-and-correct (gated KL) Student-proposed, teacher-corrected Token General Bridges teacher-student gap via interleaved sampling.
DistiLLM-2 Contrastive: Skew-FKL on teacher data + Skew-RKL on student data Mixed Token General Asymmetric losses on each data source; ICML 2025 oral.
DSKDv2 KL in dual aligned space; explicit on-policy mode Student Token Cross-tokenizer Cross-vocabulary distillation; supports both on/off-policy.
Constrained OPD KL-constrained CMDP Student Token General Hard KL constraint instead of soft penalty. Borderline OPD-RL.
AdaSwitch Adaptive on/off-policy switching Mixed Token General Switches between teacher-data and student-rollout based on divergence threshold.
Veto Logit-space geometric bridge with adaptive gradient veto Student Token General Adaptive Target Reformulation.
G-OPD / ExOPD Reverse KL + scaled reward extrapolation Student Token General Generalises OPD as KL-constrained RL; allows reward scale > 1 to "exceed" the teacher.
Fast OPD Prefix-truncated distillation reducing FLOPs Student Token (truncated) Reasoning 2× to 47× speedup via reasoning-prefix truncation.
Entropy-Aware OPD Switch between FKL and RKL based on teacher entropy Student Token Reasoning When teacher entropy high → FKL; low → RKL.
REOPOLD Mixture-based reward clipping + entropy-based dynamic sampling Student Token Reasoning "Relaxed OPD"; views OPD as policy optimisation with teacher-student log-ratio reward.
PACED Frontier curriculum at student competence boundary Student Token General Self-distill style (privileged-context / earlier-checkpoint); difficulty weighting w(p)=p(1−p).
TSD-KD Indirect (student-propose / teacher re-rank) + direct selective logit KD Mixed Token (selected) General Hybrid; partial OPD + partial preference.
SCOPE Teacher-PPL-weighted KL on incorrect rollouts; student-PPL-weighted MLE on correct Student Token Reasoning Signal-Calibrated OPD with Dual-Path Adaptive Weighting; verifier-routing.
TIP Top-50% high-entropy student tokens carry the OPD signal Student (selected) Token (filtered) Reasoning ~47% memory savings; only entropy-high student tokens trained.
HPD Reweighted log-likelihood unifying FKL + RKL Mixed (off-policy + lightweight approximate on-policy sampling) Token General Unifies KD as token-level reweighted likelihood; lightweight on-policy sampling preserves training efficiency.
TRD Trajectory-level refinement of student rollouts, then distillation Student (refined) Trajectory → Token Reasoning Argues dense per-token supervision causes prefix failure; revises problematic student predictions at the trajectory level before distilling. Generalises to on-policy self-distillation.
BRTS Token-level KL on student rollouts + auxiliary teacher-context loss Student + Best-of-N-selected teacher rollouts Token Reasoning (AIME/AMC) Selection waterfall = correctness → student-alignment → ground-truth-guided recovery; the curated teacher trajectory replaces a single high-variance teacher rollout.
OPRD Layer-wise representation alignment (deterministic; avoids Monte-Carlo KL variance over large vocab) Student Hidden-state (per-layer) Reasoning (competition math) Lifts distillation from output space into hidden-state space — "bypasses the LM head entirely". Teacher access is white-box (representations rather than logits): the first feature-based OPD.
FiRe-OPD RKL with hard trajectory filtering + soft token reweighting Student (filtered) Trajectory (hard) → Token (soft) Reasoning + Code Decouples granularity: hard filtering wins at the trajectory level, soft weighting beats hard selection at the token level. Weight w_t=(1+α·c^T_t)(1+β·c^S_t) from teacher confidence × student confusion, per-trajectory normalised. +6.25 AIME24 (Qwen3-4B, 30B teacher); multi-teacher math+code variant.
IW-OPD Per-token OPD advantage × detached, normalised prefix-importance weight Student Token (position-reweighted) Reasoning + code Diagnoses a position bias: student rollouts drift out of the teacher's distribution as they lengthen, so late tokens carry degraded supervision. Derived as the closed-form optimum of min_q D_KL(q‖π_T) s.t. D_KL(q‖π_θ) ≤ ρ, changed of measure back to π_θ. Stabilised with log-space scaling, unsigned accumulation, within-sample normalisation. No extra teacher forward passes. +6.9 AIME-2025 at step 10; gains grow as the student shrinks (+1.0/+1.2/+1.9 for 4B/1.7B/0.6B) and as the compression ratio rises (+4.0% at 1.0× → +14.9% at 6.7×). Composable with ExOPD.
DEAR Teacher–student log-prob gap advantage on a selected token subset (D ∪ E, ~36% of tokens) Student (selected) Token (filtered) Reasoning + code Direct methodological rebuttal to TIP. Entropy selectors find decisions (where to branch) but the substantive knowledge sits in evidence tokens — low-entropy, high-divergence positions where the student is confident yet wrong, structurally unreachable by any entropy criterion. Stage 2 scores non-decision tokens by hidden-state cosine similarity to decision anchors × normalised divergence. Gradient-mass coverage: 39.1% (decision-only) / 35.9% (random) / 75.8% (DEAR). Up to +2.5 pp competition math, +5.7 pp code.
ReNIO Standard token-level OPD divergence × a per-sample weight from clipped "pivotal tokens" Student (reweighted) Sample-level weight on token loss Math + code Controlled filtering shows incorrect-only training beats correct-only under both OPD (+2.59) and OPSD (+2.50), and yields longer responses with more reflection markers. Because correctness labels need full answer-bearing rollouts (defeating OPD's short-prefix cost advantage), the weight is built from prefix-computable log-ratios instead. Evaluated in both external-teacher OPD and teacher-free OPSD (teacher = the model's own initial parameters) modes; the OPSD tables are the stronger half.
PG-OPD Unchanged reverse KL; the contribution is rollout budgeting Student (prefix-screened) Token Math reasoning All K candidates decode to a fixed prefix; teacher–student top-k overlap over the first R probe tokens scores each; only high scorers (plus a guaranteed per-prompt best) continue to full length, while prefix tokens of all candidates still supervise. Up to +4.80 avg over OPD (59.40→64.20) at 2.46× wall-clock; beats PRUNE-OPD.
SEAD Zone-dependent: skip (both low-entropy) / RKL (teacher confident, student uncertain) / FKL (teacher uncertain) Student Token (zoned) + prompt (curriculum) Math reasoning Three scales of competence-adaptivity. Token: joint top-k entropy assigns ~50% of positions zero gradient, ~40% RKL to sharpen, ~10% FKL to preserve multi-path diversity. Temporal: α cosine-anneals 0.8→0.0, a continuous exploration→refinement transition. Prompt: a competence-gated curriculum admits prompts as the student's measured pass rate allows. 64.0 avg vs 59.2 vanilla OPD; a 2³ factorial shows curriculum alone is the single strongest factor (+4.20) while zones + annealing are super-additive.
DOPD Four-way routed: Top-K RKL→teacher / stop-grad self-anchor / full-vocab JS→teacher / RKL→privileged student Student (non-privileged rollout) Token (routed) Reasoning + VLM The privileged-teacher trick raises the ceiling but conflates two gaps: the transferable capability gap the student is meant to close, and the information-asymmetry gap it can only mimic. Uniform distillation therefore teaches privileged shortcuts and collapses entropy. A per-token gap A = |log π_T − log π_S| measured under identical privileged conditions selects the regime. Qwen3-8B→1.7B avg 51.4 vs 43.9 vanilla OPD, recovering 89.8% of the teacher–student gap.
MOPD (Xiaomi) Per-token reverse KL; PG form with clipped  = sg[log π_teacher − log π_student], or top-k (k=64) with a bias-correction term Student Token Multi-domain (Math, IF, SWE, Code, Tool Use) The canonical inter-stage consolidation recipe: one shared SFT checkpoint → fully parallel per-domain RL experts → student re-initialised from that same SFT checkpoint and distilled from domain-routed teachers on its own rollouts. Teachers run as standalone prefill services so teacher cost hides behind student sampling. Normalised score 0.9373 vs 0.8818 (Mix-RL) / 0.8241 (off-policy finetune) / 0.8574 (param-merge). Same-origin teachers are essential — substituting the stronger but distant Qwen3-235B-A22B collapses it to 0.60 (PG) / −1.19 (top-k) with divergence at step 18. Iter-2 reaches 0.986.
Blockwise Drift Gating Existing OPD loss × detached gate g = exp(−τ|s|) from block-aggregated old↔current drift Student (rollout-reuse) Block (64-token / newline span) Math reasoning Targets the PPO-style setting where one rollout is reused across epochs, so π_θ drifts from the π_old that generated it. Teacher targets, support and rollout policy are all untouched. LSM+Block64 is the best trained student at 53.3 avg (vs LSM 49.8, base 40.1, teacher 65.2).
Direct-OPD Teacher log-ratio r_t = log π_T − log π_T_ref (the teacher's implicit reward), Rao–Blackwellised over top-k, with adaptive α anchoring to the student's init Student Token Reasoning Weak-to-strong: run cheap RL on a small model, then transfer what that run learned to a bigger target. Vanilla OPD toward a weak post-RL teacher actively degrades a stronger student (R1-Distill-7B 56.7→~50); reading the pre-RL→post-RL shift instead improves it. The objective is itself KL-regularised RL anchored at the student's own initialisation. Qwen3-1.7B AIME24 48.3→58.3 (~4 h, 8×A100, vs Polaris direct RL at 32×A100 for a week); shifts compose (48.3→58.3→63.8). Transfer works without rising teacher–student top-k overlap, so it is not progressive imitation.
OPD² (Delta Distillation) Delta reward R^Δ = log π*(y_t) − log π*_base(y_t), both advantages top-k-centred, with a sign-agreement gate Student Token Math + science + code Isolates what post-training added to the teacher rather than the teacher's absolute distribution, so the student is not pulled toward capabilities it already shares with the base model. The sign-agreement gate zeroes updates where the delta and ordinary OPD advantages disagree, fixing the convergence point. Implemented on TRL GRPOTrainer, 100 steps. Qwen3-1.7B math avg 34.8→54.6 (vs 51.0 OPD, 51.4 ExOPD); Qwen3-8B AIME24 76.2.
TOP-D Probability-space teacher–student interpolation collapsing to r̃ = log(αρ + 1 − α); GRPO-style clipped surrogate with token-level normalisation Student (rollouts reused across mini-batch epochs) Token Math reasoning The smoothed reward is strictly lower-bounded, so gradient variance is bounded (Thm 4.2) instead of exploding as π*→0 — at zero extra compute over standard OPD. Comes with a global convergence bound and a monotonic-improvement bound. Qwen3-8B-Base ← Qwen3-30B-A3B: AIME24 avg@32 50.42 vs 24.58 standard OPD, AIME25 +10.73, AIME26 +18.64; vs GRPO 30.10 / DAPO 32.92.
COPD Log-likelihood contrast A_t = ℓ^LT − ℓ^HT between two teacher scorings of the same student tokens, clipped and used as a PPO-style advantage Student Token Multimodal reasoning Instead of matching a distribution, it asks the teacher which of two reasoning modes the token belongs to. No accuracy reward, no verifier, no length penalty. vs ExOPD on Qwen3-VL-8B→2B: +2.1 pp with 57.0% shorter responses (Acc@1K 26.7→64.4) at 60 GPU-h vs 132 (OPD) / 150 (ExOPD). The COPSD variant swaps in a frozen snapshot of the student itself: +3.5 pp, −63.8% length.
ShortOPD Generalized JSD (α=0.5) over top-100 logits + aggregated tail mass Student (adaptive horizon) Token Compression (post-pruning recovery) Pruning Qwen3-4B-Instruct by 4/36 layers collapses greedy GSM8K 88.1→49.0 but pass@64 recovers to 91.2 — correct trajectories are demoted, not erased, so recovery needs on-policy states and dense targets with the frozen pre-compression parent as teacher (no labels, verifier, or external teacher). At 25% pruning 55–75% of early rollouts end in repetitive suffixes carrying ~35× less signal, so EMAs of repetition / truncation / effective length adapt the per-step horizon. Avg over 8 tasks 5.71→48.46 = 64.5% of the dense teacher, +17.9 over the best off-policy baseline, 8.5 h vs 35.9 h.
W2S-OPD Per-token reverse KL toward a synthesised teacher softmax(z_base + α(z⁺ − z⁻)), estimated on the teacher's top-K support Student Token Math + code Weak-to-strong: every supervision source is smaller than the student. Because the weak pair enters only through its logit difference, what the two weak models agree on (their shared limited ability) cancels, and only the direction along which m⁺ improves over m⁻ transfers; re-anchoring that direction on the student's own base keeps the target inside a distribution the student already realises, with α bounding how far it moves. Eq. 2 is an exponential tilt of the student's base, so the proxy teacher is the closed-form argmax of E_q[r] − (1/α)·KL(q‖π_base) — a trust region centred on the student, not on any weak model. Three contrast sources are interchangeable (post-RL vs pre-RL expert · Qwen3-4B vs 0.6B base · one model under correct vs wrong hints), and several deltas sum on the shared anchor for multi-teacher merging. Qwen3-8B math avg 17.0→51.8 (vs 46.5 OPD, 45.7 SFT), surpassing the 4B expert itself (48.8); +6.0 pp from two off-the-shelf base models both weaker than the student; +1.4 pp from a single model under contrastive hints. OOD GPQA-Diamond 38.9→56.5, and IFBench improves where OPD drops below the base. A top-1% highest-Δ token analysis over Schoenfeld episodes separates the sources: the post-RL and hint contrasts reinforce Plan/Monitor, the scale contrast Analyze/Implement.
TOPD (masked dLLM) Token-level reverse KL via a sampled-token score-function estimator, on trace-aligned decisions only Student (own denoising trajectory) Token (trace-aligned) Reasoning (diffusion LMs) Random-mask supervision — the default in dLLM RL/SFT pipelines — can reveal later answer tokens while hiding earlier reasoning ones, creating backward-reconstruction states the student never visits at inference. TOPD instead rolls out the real low-confidence-remasking decoder and keeps only commitments that survive into the final answer. SDAR-4B-Chat ← TraDo-8B-Instruct: MATH500 70.2→75.9, matching TraceRL-trained TraDo-4B with 4× fewer rollout rounds. Ablations: on-policy 75.2 > off-policy 74.5 > semi-AR SFT 73.5; trace-aligned 75.2 > random-mask 74.3; RKL > JSD > FKL.

📝 Strictness notes

  • BRTS — ⚠️ Partially dilutes C1: the primary student-context leg is strict OPD (student trains on its own rollouts), but the auxiliary teacher-context branch supervises on teacher-generated (off-policy) trajectories. Listed because the student-context leg is the core objective and the teacher branch only stabilises it.

  • OPRD — ⚠️ Not logit-based: C1 ✓ / C2 ✓ on student rollouts, but supervision is feature/representation-space (hidden states across layers), not next-token logits. Listed in White-Box because teacher access is white-box; flagged here because the "feature" supervision signal sits outside the section's default logit-matching form.

  • Direct-OPD — ⚠️ Two departures from the section default. The supervision is a teacher log-ratio (implicit reward), not the teacher distribution, so it is arguably an OPD/RL hybrid; and the teacher is smaller than the student (weak-to-strong), which inverts the usual strong-to-weak setting. C1 ✓ / C2 ✓ — the teacher is still queried on the student's own visited prefixes.

  • OPD² — ⚠️ Stronger access assumption than ordinary white-box OPD: three models must be loaded (student, teacher, and the teacher's pre-post-training base checkpoint), which is not available for most released teachers.

  • COPD — ⚠️ C2 is satisfied at token granularity (teacher log-probs on student tokens) but the loss is not a KL/distribution-matching objective — the teacher signal is converted into an RL advantage. Listed in White-Box rather than OPD-RL Hybrids because there is no reward model or verifier anywhere in the objective. Its COPSD variant is squarely OPSD.

  • TOP-D — ⚠️ C1 slightly relaxed: the internal trust-region iterations deliberately reuse rollouts across mini-batch epochs, so updates are near-on-policy. The paper's own "w/o off-policy" ablation is the strict variant.

  • Blockwise Drift Gating — ⚠️ Authors describe it as "a preliminary empirical study": one student, one teacher, one dataset, no repeated seeds, and AIME sets with tiny sample counts — a 1.7-point pass@8 delta on 4 benchmarks is plausibly within noise. Also a pure loss-weighting heuristic on an existing OPD loss, not a new supervision mechanism.

  • ShortOPD — teacher is the student's own uncompressed parent, so it sits between the compression slot and OPSD. C1 ✓ / C2 ✓ otherwise textbook.

  • W2S-OPD — ⚠️ Inverts the section default: every supervision source is smaller than the student (weak-to-strong), so "larger external teacher" describes only the synthesised proxy, not any real model on disk. Three frozen models must be loaded (the student's own base as anchor + the contrast pair), though only the post-RL instantiation needs a post-training checkpoint pair; the scale and hint variants use two off-the-shelf base models or a single model under two prompts. C1 ✓ / C2 ✓ — the delta is composed in logit space and supervision remains a full top-K distribution matched by reverse KL, so the objective stays distribution-matching rather than an OPD/RL hybrid.

  • ⚠️ Acronym collisions. This field has reused several abbreviations for unrelated work; always resolve by arXiv ID:

    • TOPD = Trace-Based OPD for masked dLLMs (2607.16872) · Near-Future-Guidance trajectory OPD (2606.00305) · Truncated OPD (2605.31490).
    • MOPD = Multi-Teacher OPD, Xiaomi (2606.30406) · Multi-Rollout OPD, Microsoft/CMU/Purdue (2605.12652) · and generically for multi-teacher consolidation in most 2026 production reports.
    • COPSD = Crosslingual OPSD (2605.09548) · Constitutional On-Policy Safe Distillation (2606.03089).
    • D-OPSD / d-OPSD / dOPSD = step-distilled image diffusion (2605.05204) · dLLM self-future (2606.18195) · dLLM peek-ahead (2607.04428).
    • PBSD = Preference-Based Self-Distillation (2605.05040, OPSD) · Posterior-Bayesian Self-Distillation (2606.09348, Agent) — unrelated mechanisms, same acronym.
    • COPD (2607.19046, SJTU/Qwen — contrastive light-vs-heavy-thinking prefixes) and CoPD (2604.27083, JD.COM — co-evolving sibling branches) differ only in capitalisation.
    • TrOPD (2606.01249, Samsung/Oxford/PKU) and TOP-D (2607.04751, Microsoft/HKUST-GZ) are different papers with different mechanisms — verified separately.
  • Teacher decay is a recurring, independently-rediscovered phenomenon. Supervision quality degrades as the student's prefix lengthens, named separately as Off-Policy Teacher Decay (ESR), Supervision Fidelity Decay (LGR), local teachability collapse (2605.13643) and depth-inverted discriminability (TurnOPD). KAT adds the sharper warning that low KL is not evidence of health — the teacher may simply be agreeing with a corrupted prefix. The prefix-truncation family (Fast OPD, Prune-OPD, PG-OPD, ESR, ADWIN, Truncated OPD) is the practical response.

  • Whether incorrect or correct rollouts carry the signal is genuinely contested. ReNIO and Apple's diagnostic find supervision is better aligned on incorrect rollouts; Yonsei's compaction study finds OPSD mainly compresses already-correct traces and barely repairs failures. Both were verified by full reads; the disagreement is real, not a listing error.

  • TOPD / diffusion-LM entries — C1 ✓ (the student rolls out its own denoising trajectory with the real inference-time decoder) and C2 ✓ at token level, but "trajectory" means a denoising path over masked positions rather than a left-to-right generation, so per-step semantics differ from the autoregressive entries.

  • REOPD (2608.11698) and CROP (2608.13387) — ⚠️ both were initially filed as OPD-RL hybrids and moved here after a full-text read. Both wear a PPO-shaped surrogate, but in neither case does any signal outside the teacher's own log-probabilities enter the objective: REOPD's advantage traces entirely through log π_θ / log π_T / log π_ref and the paper states it "requires no verifier, reward model, value model, or extra rollout beyond standard OPD"; CROP is a hard token mask over the unchanged clipped OPD surrogate. Same filing logic as G-OPD (2602.12125), whose reward extrapolation also lives here.

  • OPTD (2608.02942) — ⚠️ diffusion-LM caveat, as already applied to TOPD (2607.16872): the supervised "trajectory" is a denoising path over masked positions, not left-to-right generation.

  • OPD² multilingual (2608.05802) — ⚠️ shares both the OPD² name and the same GitHub repo (naver-ai/opd2) as the existing 2607.15161 entry. Companion second paper, not a replacement; the existing note about requiring the teacher's pre-post-training base checkpoint applies here too.

  • New acronym collisions this window — SPOT: 2608.04419 (sparse probing, White-Box) vs. the listed 2603.01683 ("Surgical Post-Training", Black-Box). TA-OPD: 2608.14728 (Tail-Aware, HuipengHuang/TA-OPD) vs. the listed 2605.26844 (token Teachability, wyy-code/TA-OPD) — same acronym, different repo, different mechanism. REOPD (2608.11698) vs. the listed ReOPD (2607.04763) and REOPOLD (2603.11137). Resolve all by arXiv ID.


🎭 OPD with Black-Box / Outcome-Based Teachers

When the teacher is API-only (no logits), OPD uses scalar rewards, verbal scores, preferences, or adversarial discriminators — all evaluated on student rollouts. Entries that turned out to use static teacher data only (Lion, SuperCorrect, DAIL, SODA) are excluded from this list.

Resource 🌟 Stars Date Org Paper Link Title / Notes
ORPO-Distill Paper 2025.09 Industrial arXiv 2509.25100 ORPO-Distill
LMOps /gad Stars 2025.11 Microsoft Research arXiv 2511.10643 · project GAD — Black-Box OPD
OVD Paper 2026.01 HKU / Huawei arXiv 2601.21968 OVD (On-policy Verbal Distillation) — project page OVD.github.io 404s
SPoT Stars 2026.03 Visual-AI arXiv 2603.01683 SPOT: Surgical Post-Training — black-box oracle edits student failures into proximal rollouts
SODA Paper 2026.04 Academic arXiv 2604.03873 SODA — Semi On-Policy Black-Box Distillation
ROPD Stars 2026.05 NUS / USTC / Tencent arXiv 2605.07396 ROPD — Rubric-based On-Policy Distillation; induces prompt-specific rubrics from teacher–student contrasts, then scores student rollouts by those rubrics (logit-free / black-box); up to 10× sample efficiency
PRISM Stars 2026.04 HKUST(GZ) / Tsinghua / NTU / RUC / USTC / UCAS arXiv 2604.28123 PRISM — Pre-alignment via Black-Box OPD for Multimodal RL; an adversarial OPD stage inserted between SFT and RLVR, scoring student rollouts with a Mixture-of-Experts discriminator (separate perception and reasoning experts, Bradley–Terry loss). +4.4 / +6.0 avg over SFT→RLVR at 4B / 8B
OmniOPD Paper 2026.06 Meta AI arXiv 2606.01476 OmniOPD — Logit-Free OPD via Speculative Verification; an entropy-driven scheduler picks uncertain chunks of the student’s rollout, a black-box teacher generates Monte-Carlo continuations, and they are scored by semantic similarity rather than logits. Beats white-box OPD with the same teacher family (69.08 vs 64.16)
ExpRL Stars 2026.06 Stanford / CMU arXiv 2606.17024 ExpRL — Exploratory RL for LLM Mid-Training; an LLM judge scores the student’s own rollouts and prefixes against a hidden reference solution under a fixed rubric, giving dense process-level reward before sparse RL ⚠️ see strictness note
SOPD Paper 2026.08 Nanjing Univ. (SKL) / XingYun Lab / UCAS / Fudan arXiv 2608.16333 Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning — at each step the teacher substitutes its own step for the student's, and the student is trained on the resulting relabelled trace; a step-granularity knob interpolates between pure OPD and pure SFT. Filed Black-Box because the teacher exposes no distribution: "the teacher supplies generated text only; SOPD requires neither teacher logits nor gradients". ⚠️ DAgger-style relabelling rather than scoring of the student's own continuation — see strictness notes. AIME24/25, HMMT25-Feb/Nov, ALFWorld

📋 Click to view technical details

Method Feedback Signal Data Granularity Domain Notes
ORPO-Distill Student-Generated Outputs (SGO) + ORPO contrastive Mixed (student-generated negatives, teacher positives) Sequence Cross-architecture "Mixed-policy strategy utilizing student-generated outputs"; NeurIPS 2025 WS.
GAD (Generative Adversarial Distillation) Discriminator (on-policy reward model) Student Sequence General A trained discriminator distinguishes student outputs from teacher (e.g. GPT-5) responses; minimax game makes the discriminator co-evolve into an on-policy reward model. Qwen2.5-14B student becomes comparable to GPT-5-Chat on LMSYS.
OVD Verbal scores (0–9) on student trajectories Student Sequence General Replaces token-level logit matching with verbal scoring; +25.7% over baselines.
SPOT Black-box Oracle step edits + BCE reward objective Student rollouts, Oracle-rectified Step / sequence Math reasoning Minimal edits keep samples proximal to the student distribution, targeting reasoning gains with knowledge retention.
SODA DPO: teacher responses as preferred vs. base student (q₀) zero-shot responses as rejected Mixed Sequence Cross-architecture "Semi on-policy" paradigm: captures student-specific inferior behaviors from a one-time static snapshot of q₀, eliminating the need for dynamic rollouts or adversarial training. 10× faster and 27% less peak GPU memory than GAD. Outperforms GAD on 15/16 benchmarks
ROPD Prompt-specific rubric scores (Rubricator + Verifier) Student Sequence (rubric-weighted reward) General Black-box-compatible alternative to logit OPD: a Rubricator contrasts teacher vs. student responses to induce prompt-specific rubrics, a Verifier scores each student rollout into a weighted pass-rate reward. Needs only teacher-generated responses, no logits.

📝 Strictness notes

  • ExpRL — ⚠️ The paper presents itself as RL mid-training, not distillation, and explicitly benchmarks against (and beats) a true OPSD baseline. It qualifies here only under the black-box reading: an LLM judge scores the student's own rollouts against a hidden reference under a fixed rubric. No teacher logits and no KL term anywhere.

  • OmniOPD — the teacher may be fully API-only; supervision is a continuous semantic-similarity score over Monte-Carlo teacher continuations at student-chosen chunks, not a distribution. Notable for outperforming white-box OPD with the same teacher family.

  • Excluded from this section after a full read: ZPPO (NVIDIA) — its own subtitle, "Teacher in Prompts, Not Gradients", states the disqualifier exactly: the teacher rewrites the prompt, the student resamples, and training is plain GRPO on binary reward. The teacher never scores a student token. It is the cleanest illustration of where C2 draws the line.

  • SOPD (2608.16333) — ⚠️ partially satisfies C2 in this section's own sense: the teacher exposes no distribution at all ("the teacher supplies generated text only; SOPD requires neither teacher logits nor gradients"), and at each step it substitutes its own step for the student's rather than scoring the student's continuation. That is DAgger-style relabelling; listed here because the relabelled states are still the student's own, and the step-granularity knob explicitly interpolates toward pure OPD.


♻️ Self-Distillation with Privileged Context — OPSD

Same model = teacher = student, but the teacher is conditioned on something the student doesn't see (verified trace, ground-truth answer, "be concise" prefix, longer context, document, …). The gap exists because of the conditioning, not weights.

Several entries previously listed here turned out on verification to use static teacher data or a fixed self-rewritten dataset rather than student rollouts; those have been excluded. SPIN was reclassified to Iterative Self-Bootstrapping.

Resource 🌟 Stars Date Org Paper Link Title / Notes
OPSD Stars 2026.01 UCLA / Meta FAIR arXiv 2601.18734 · blog OPSD — Self-Distilled Reasoner
Self-Distillation Stars 2026.01 MIT / ETH arXiv 2601.19897 SDFT-Continual
mtp-lm Stars 2026.02 UMD / LLNL arXiv 2602.06019 MTP Self-Distill
LMOps /opcd Stars 2026.02 Microsoft Research arXiv 2602.12275 OPCD — On-Policy Context Distillation
GATES Paper 2026.02 UMD arXiv 2602.20574 GATES (Self-Distillation under Privileged Context)
EMPO² Paper 2026.02 Microsoft Research arXiv 2602.23008 · code · blog EMPO² — memory-tip-conditioned online self-distillation for exploratory LLM agents (ICLR 2026; cross-listed into Agent)
CRISP_Reasoning_Compression Stars 2026.03 LinkedIn arXiv 2603.05433 OPSDC / CRISP
LMOps /oel Stars 2026.03 Microsoft Research arXiv 2603.16856 OEL — Online Experiential Learning
self-distillation-analysis Stars 2026.03 MSR / KAIST / SNU arXiv 2603.24472 Why Does Self-Distillation (Sometimes) Degrade Reasoning? — diagnostic study of OPSD failure modes
ml-ssd Stars 2026.04 Apple MLR arXiv 2604.01193 Apple — Embarrassingly Simple Self-Distillation
Skill-SD Paper 2026.04 UCAS / CUHK / USTC / vivo AI Lab arXiv 2604.10674 Skill-SD — skill-conditioned OPSD for multi-turn LLM agents
SD-Zero Paper 2026.04 Princeton / Toronto / CMU arXiv 2604.12002 SD-Zero — Self-Revision turns binary rewards into dense supervision
π-Play Paper 2026.04 CASIA / UCAS / Meituan arXiv 2604.14054 π-Play — multi-agent self-play turns the question-construction path into privileged context for OPSD on search agents
OPSDL Paper 2026.04 Baidu arXiv 2604.17535 OPSDL (Long-Context Self-Distillation)
MSD Paper 2026.05 Tongji / Shanghai AI Lab arXiv 2605.02971 MSD — multilingual safety OPSD; teacher conditioned on English query translation + CoT instruction; DPSW weights safety-critical tokens
COPSD Stars 2026.05 LMU Munich / MCML arXiv 2605.09548 COPSD — crosslingual OPSD; teacher sees English problem translation + reference solution, student rolls out in low-resource language (17 African languages)
SGSD Stars 2026.05 THU arXiv 2605.28791 SGSD — Skill-Conditional Gated SD
CODE Stars 2026.05 USTC arXiv 2605.28303 CODE — OPSD on Knowledge Editing + Casual Editing
SSOPD Paper 2026.05 THU / Beihang arXiv 2605.17497 SSOPD — Self-Supervised OPSD; privileged context is the model's own shortest correct completion within a GRPO group (no external traces), distilled into prefixes of the longest wrong completion
RLCSD Stars 2026.06 THU (BPM) / Alibaba Tongyi arXiv 2606.11709 RLCSD — Contrastive OPSD; cancels privilege-induced style drift by contrasting the teacher–student gap under a correct hint vs. a wrong hint; verl-based
d-OPSD Stars 2026.06 THU / TUM / NTU / UT Austin arXiv 2606.18195 d-OPSD — first OPSD for diffusion LLMs; self-generated answers as suffix conditioning ("self future-experience"); step-level (not token-level) divergence aligned to the denoising process
CaOPD Stars 2026.04 Salesforce AI Research arXiv 2604.16830 CaOPD — The Illusion of Certainty; proves privileged conditioning makes the teacher's confidence a non-identifiable and upward-biased target, so OPD reliably buys accuracy at the cost of severe overconfidence; fixes it by rewriting confidence targets to the free empirical rollout success rate
Vision-OPD Stars 2026.05 ISCAS / UCAS / Xiaohongshu arXiv 2605.18740 Vision-OPD — regional-to-global OPSD for MLLM fine detail; teacher sees a 2×-upscaled evidence crop, student sees the full image. 9B beats Gemini-3.1-Pro on the fine-grained suite using 6.2K fully synthetic triplets, no GT labels or verifier
D-OPSD Stars 2026.05 HKUST / Z-Image Team Alibaba / UCSD / CUHK arXiv 2605.05204 D-OPSD — OPSD for continuously tuning step-distilled diffusion models; exploits that LLM/VLM-encoder T2I models inherit in-context ability, so feeding the encoder the target image is free privileged context
PW-OPSD Stars 2026.05 SaFo Lab / UW–Madison arXiv 2605.21606 PW-OPSD — When Are Teacher Tokens Reliable?; a branch-viability diagnostic shows position separates genuinely-uncertain from merely-diverse teacher tokens (AUROC 0.83) where every local uncertainty measure fails (≤0.57)
OPSD Predictive Law Stars 2026.05 Tufa Labs, Zürich arXiv 2605.30070 A Predictive Law for OPSD From World Feedback — the pre-training student–self-teacher accuracy gap linearly predicts final OPSD gain (R² 0.949 / 0.996), so privileged-context designs can be screened before training (ICML RLxF 2026)
SDSD Diversity Paper 2026.06 Mila / Univ. de Montréal / FAIR at Meta arXiv 2606.26091 OPSD with Sampled Demonstrations Reduces Output Diversity — proves the SDSD optimum tilts the base policy by pointwise conditional mutual information rather than reward, so it amplifies dominant modes; pass@1 up, pass@k flat
PHF Paper 2026.06 HKUST(GZ) / NUAA / NUDT arXiv 2606.29340 PHF — Privileged Hidden Flow; adds a residual-stream transition-geometry channel (direction + Gram-matrix CKA) on top of the standard OPSD output loss, with proven invariance to per-trajectory offsets and rescaling
Visual-OPSD Stars 2026.06 XJTU (MOE KLINNS) / SYSU arXiv 2606.18974 Visual-OPSD — shows a unified multimodal model's rendered "visual thoughts" matter as a generation pathway, not as pixels, then distills that pathway into a text-only student: +3.40 pp over its own teacher at 14.3× speedup
Denser ≠ Better Stars 2026.07 HKISI CAS / CASIA / UCAS / NJUST arXiv 2607.01763 Denser ≠ Better: Limits of OPSD for Continual Post-Training — SDPO specialises harder than GRPO but forgets far more (ToolUse −80% vs GRPO +18%); dense self-distillation is a specialisation accelerator, not a continual-learning stabiliser
Purified OPSD Paper 2026.07 ZJU / Tongyi Lab Alibaba / HUST / Jilin arXiv 2607.02234 Purified OPSD — Without Losing How to Think; decomposes the teacher update via a reference-only teacher and finds the reference-memorisation component dominates while the useful component is actively opposed (cos ≈ −0.95); re