跳到正文
Dongxi 东锡 NLP· @dongxi_nlp · X·· 4 小时前AI 评分32
AI 导读

推荐阅读! ExploreNet: Learning Where to Explore in Diffusion GRPO 我们推出 ExploreNet🔍,用于扩散RL中的可学习探索。 ExploreNet 学习以当前状态为条件的自适应探索分布,以 rollout 的多样性作为奖励。 ➡️ 在扩散 GRPO 中实现更快、更有针对性的学习。

正文

推荐阅读!

ExploreNet: Learning Where to Explore in Diffusion GRPO

引用Stella Li @ COLM@StellaLisy
We introduce ExploreNet🔍 for learnable exploration in diffusion RL. ExploreNet learns an adaptive exploration distribution conditioned on the current state, rewarded by the diversity of the rollouts. ➡️ Faster, more targeted learning in diffusion GRPO.
在 X 查看被引用的帖子

来源:Dongxi 东锡 NLP · x.com