Kolibri-1 无需微调自主玩 Breakout,每次决策推理延迟低于 25ms
Less Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move.
Kolibri-1 在无需微调的情况下自主玩 Breakout,每次决策推理延迟约 25ms。该实验测试模型在少量约束下输出结构化动作概率的能力,全程不生成文本,仅输出四个动作,权重已开源。
| Got Kolibri-1 to play Breakout completely on its own, no fine-tuning. The more we explore u/Aleph__Alpha’s Kolibri the more it get's exciting and its potential. Less talk. More Breakout is one such experiment to see how good the model is at structured output given a few constraints. We especially optimized the inference for action probabilities: around 25 ms inference per decision. Four moves. No generated text. One shared game. Open weights. New possibilities. Watch it play: https://tesseracted.com/kolibri-1-chat/gameplay/breakout/ [link] [留言] |
来源:r/LocalLLaMA · reddit.com