CoRT:面向词元级评分规则引导策略优化的反事实回放方法
英文摘要
A research paper introduces CoRT, a method that employs counterfactual replay to enable token-level, rubric-guided policy optimization. The approach aims to align language model outputs with granular, token-wise scoring rubrics. No implementation details or benchmark results were provided in the shared content.
中文摘要
一篇研究论文提出了 CoRT 方法,该方法利用反事实回放实现词元级、评分规则引导的策略优化。其目标是让语言模型的输出对齐细粒度的词元级评分标准。分享的内容中未提供实现细节或基准测试结果。
关键要点
CoRT uses counterfactual replay for policy optimization.
CoRT 利用反事实回放进行策略优化。
The optimization is guided by token-level rubrics.
优化过程由词元级评分规则引导。
Only the paper title and a link were shared, no further technical details.
仅分享了论文标题和链接,无进一步技术细节。