Blockway 发布 Agens Volundr 32B Preview,72 层中仅 18 层保留 KV cache
Agens Volundr 32B Preview: our small team's first model on our own hybrid architecture. Only 18 of 72 layers keep a KV cache (Apache-2.0)
香港小团队 Blockway 发布 Agens Volundr 32B Preview,这是其自研混合架构上的首个模型,采用 Apache-2.0 许可。架构共 72 层,其中 54 层为 KDA 线性注意力、17 层为自研压缩稀疏注意力 BCSA、1 层为全注意力,因此只有 18 层保留 KV cache,上下文窗口 262K。
| Hi r/LocalLLaMA. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front. WHY WE BUILT IT Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one. ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)
So only 18 of 72 layers keep a KV cache. Context window: 262K. SPEED (single user, our sglang build)
BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)
KNOWN LIMITATIONS (please read before trying)
RUN IT docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200) The full launch command is in the model card. LINKS
Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments. submitted by /u/ComfortableKindly507[link] [留言] |
来源:r/LocalLLaMA · reddit.com