我跑了 1200+ 次 Blender 建模测试,横跨多个 LLM 与智能体框架,并做成可投票的盲测竞技场
I ran 1,200+ Blender modeling runs across LLMs and agent harnesses and made them votable
一位开发者搭建了让 AI 智能体在 Blender 中建模的盲测 A/B 竞技场,用户可投票比较结果,目前已积累 1200+ 次运行。测试覆盖 pi、opencode、omp、codex、Claude Code 和 dsh,并对比直接写脚本与使用 MCP 的智能体,同时考察不同推理等级的成本与时间表现。新增参考图建模模式,支持在侧边栏切换文本与图像模式。
Blind A/B arena where AI agents build things in Blender and you vote on which result is better: https://render-arena.izolight.xyz
Each agent gets a prompt that describes its environment and the rules, plus a few words for what to model. It's inspired by minebench.ai (initial prompts are borrowed from there), and I wanted to see whether the same progression across models shows up.
What I think few arenas cover is the harness, not just the model. I ran pi, opencode, omp, codex, Claude Code and dsh, and compared agents that write scripts straight into Blender with ones that have an MCP. I also covered the reasoning levels, mainly to find cost and time sweet spots.
It has 1,200+ runs, but not every combination for every prompt, because that would get expensive. You can submit your own runs if you want to help fill gaps.
I just added a second mode where the agent gets a reference image and has to model it as accurately as it can. You switch between text and image mode in the sidebar. It has one image and few runs so far, and will grow.
Votes are what make the rankings mean anything, so a few minutes of voting helps a lot. Feedback on the method is welcome.
submitted by /u/Izolight
[link] [留言]
来源:r/LocalLLaMA · reddit.com