跳到正文
r/MachineLearning· /u/Yossarian_1234·· 4 小时前AI 评分32

MA-BC:多目标模仿学习中的分差合并方法,具备可证明的样本复杂度上下界

Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation [R]

AI 导读

针对"如何向目标各异的专家学习"这一问题,研究者提出 MA-BC 方法:仅在专家演示中动作不冲突的部分进行数据池化,从而避免全量池化丢失权衡、分别学习又无法共享数据的问题。该工作给出了样本复杂度的上界与下界,作者为 Ziyad Sheebaelhamd、Luca Viano、Volkan Cevher、Claire Vernade。

正文
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation [R]

https://preview.redd.it/i2c0cdkg04uh1.png?width=2532&format=png&auto=webp&s=fdfed9fe7ed2cef8e45791110a5152ff163a0ce8

TLDR: The question we answer: how do you learn from experts with different objectives? Pooling all their data can lose their trade-offs; learning from each expert separately misses opportunities to share data. MA-BC pools demonstrations where observed actions don’t disagree, with upper and lower bounds on sample complexity.
Authors: Ziyad Sheebaelhamd, Luca Viano, Volkan Cevher, Claire Vernade

submitted by /u/Yossarian_1234
[link] [留言]

来源:r/MachineLearning · reddit.com