r/MachineLearning· /u/Yossarian_1234·· 4 小时前AI 评分32
MA-BC:多目标模仿学习中的分差合并方法,具备可证明的样本复杂度上下界
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation [R]
AI 导读
针对"如何向目标各异的专家学习"这一问题,研究者提出 MA-BC 方法:仅在专家演示中动作不冲突的部分进行数据池化,从而避免全量池化丢失权衡、分别学习又无法共享数据的问题。该工作给出了样本复杂度的上界与下界,作者为 Ziyad Sheebaelhamd、Luca Viano、Volkan Cevher、Claire Vernade。
正文
| TLDR: The question we answer: how do you learn from experts with different objectives? Pooling all their data can lose their trade-offs; learning from each expert separately misses opportunities to share data. MA-BC pools demonstrations where observed actions don’t disagree, with upper and lower bounds on sample complexity. [link] [留言] |
来源:r/MachineLearning · reddit.com