如何在 NVIDIA GB300 NVL72 上部署 Qwen3.8-2.4T-A95B 并配置推理
Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
NVIDIA 开发者博客介绍如何在 GB300 NVL72 上部署阿里开源的 Qwen3.8-2.4T-A95B(Qwen3.8-Max),并支持可配置推理。该模型总参数 2.4T、每 token 激活 95B,采用细粒度 MoE 架构与全注意力加线性注意力的混合设计,上下文窗口最高 100 万 token。
原文给出在 NVIDIA GB300 NVL72 上部署 Qwen3.8-2.4T-A95B 并配置推理的具体做法,可迁移到同类大 MoE 模型的部署场景。
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open...
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open ecosystem. It has 2.4T total parameters with 95B activated per token. It’s a fine-grained mixture of experts (MoE) architecture with a hybrid of full and linear attention, a context window of up to one million tokens, and an output length of up to…
来源:NVIDIA Developer Blog · developer.nvidia.com