你好,我是 Palind,这里是我的博客~ 你可以在我的中查看我的学历、荣誉、爱好等。
You’re very welcome to check out my blog! You can find my education, honors, hobbies, and more in my
我正在学习 LLM 和 agent。
你好,我是 Palind,这里是我的博客~ 你可以在我的中查看我的学历、荣誉、爱好等。
You’re very welcome to check out my blog! You can find my education, honors, hobbies, and more in my
我正在学习 LLM 和 agent。
还真是,《score-centering的前世今生》 与 《OPENAI的神秘RL算法只是个STE的近似吗?解密OAI的RL 后训练算法》 都是这么解释的。☺️
上一次的学习笔记:Claude Code 学习笔记 | Palind’s Blog。
这一次主要围绕多 agent 的机制。
主要内容:
DPO 极简笔记,方便回顾。