ニュース
Pass the Baton: Trajectory-Relayed On-Policy Distillation
arXiv cs.AI · 公開日 · 読了3分
30秒で要点
- 何が起きたか
- Relay-OPD improves student model training by having a teacher model take over when the student commits to wrong reasoning directions, then resuming student training on corrected trajectories.
- なぜ重要か
- Matters for engineers training smaller language models on reasoning tasks who want better performance with less compute spent on misdirected generations.
- 注意点
- Method requires detecting failure points and managing teacher intervention budget; unclear how well trigger detection generalizes across different model sizes and task domains.
この要約を音声で聴く
記事より
-->
Computer Science > Computation and Language
arXiv:2607.26057v1 (cs)
[Submitted on 28 Jul 2026]
Title: Pass the Baton: Trajectory-Relayed On-Policy Distillation
Authors: Haolei Xu , Xiaowen Xu , Haiwen Hong , Zixuan Ni , Hongxing Li , Yiwen Qiu , Weiming Lu , Yongliang Shen
View a PDF of the paper titled Pass the Baton: Trajectory-Relayed On-Policy Distillation, by Haolei Xu and 7 other authors
View PDF HTML (experimental)
Abstract: On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). During training, Relay-OPD constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy. With a Qw
原文からの抜粋です。全文は配信元でお読みください。
- reasoning
The Agent Architect
1つのパターン、1つのトレードオフ、1つの本番障害事例。エージェントシステムを構築する人のための短い週刊ブリーフィング。
週1回のメール、ワンクリックで購読解除できます。アドレスはブリーフィングの送信のみに使用します。