CivBench 受控版评测 LLM 玩《文明 5》:GLM-5.3 领先 Opus-5.5,Qwen-3.8-27B 表现亮眼
A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.
CivBench 受控版让模型在《文明 5》的三个固定开局中轮换对战,每局两个文明由 LLM 担任战略家、六个由 Vox Populi AI 执行,以测试 50 到 100+ 回合后才显现后果的决策。
| A while ago, I posted here getting OSS-120B and GLM-4.6 playing full games of Civilization V. Since then, models have moved pretty far, and we wanted a better understanding about models' capabilities playing the game. Introducing the controlled version of CivBench on newer models: The controlled version of CivBench (Chen et al., 2026, extending our COLM 2026 work) We are currently testing GPT-6.1-Sol, GPT-6-Astra, etc. Feel free to suggest some models (especially interesting open-weight ones) for our next run! What is Civilization? Civilization V ($7.49 today on Steam promotion) is a turn-based strategy game where you take a civilization through hundreds of turns of expansion, science, diplomacy, war and eventually the space age. That makes it useful for testing something LLM benchmarks often struggle with: decisions whose consequences may not show up until 50 or 100+ turns later. A LLM strategist playing as Byzantine. Can Theodora rebuilt Rome? What makes this a controlled experiment? Instead of giving each model unrelated games, we rotate them through the same three fixed starts. Each game has eight civilizations: two using the tested LLM strategist and six using the standard Vox Populi AI. The LLM sets high-level strategy; Civ's existing AI handles low-level execution. Can I see how the models actually play? Yes. A few examples:
Can I play a round now? Yes. If you own the game, Vox Deorum is open source and has an installer. You can play Civilization V yourself against LLM-powered civilizations, watch a full AI-vs-AI game, or even chat with your opponents. You can also have LLMs as your teammates and work together towards a win! Guess I can't avoid an unequal treaty as a pacifist. At least I can get a bargain? Can I use local models or my existing subscriptions? Yes. Local OpenAI-compatible servers are supported, and Qwen-3.8-27B can do an excellent job. You can also use your existing Claude or Codex subscriptions. (I use them to run a ton of evaluation games! GPT-6-Luna is basically free to play. About $0.5 in API cost per player per game.) What else did you learn? Please check out our COLM 2026 paper for methodology and EMNLP 2026 paper for whether models would authorize nuclear strikes on others. I guess Civilization is just a game, don't you think so? Can we at least have a chat, please? (Sorry for sending and deleting this repeatedly. Guess I shouldn't use in-flight wifi to send a post with many pictures. I hope they go through! Please let me know if you can't see them.) submitted by /u/vox-deorum[link] [留言] |
来源:r/LocalLLaMA · reddit.com