跳到正文
r/LocalLLaMA· /u/Prestigious-Taste-63·· 14 小时前AI 评分22

从零训练 3.87B MoE(1.45B 激活)模型 Apex-2,仅用 86.5B tokens

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

AI 导读

开发者从零训练了 MoE 模型 Apex-2,总参数 3.87B、每 token 激活 1.45B,仅用 86.5B tokens 预训练,采用 Qwen3 tokenizer、16 专家 top-4、4096 上下文。

正文
I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

First of all, thank you for reading.

I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons.

Apex-2

- Architecture: Decoder-only MoE, every layer is MoE (no dense layers)

- Size: 3.87B total parameters, 1.45B active per token

- 32 layers, d_model 2048, GQA 16Q/4KV, 16 experts, top-4

- Context: 4096

- Tokenizer: Qwen3 (151k)

- Hugging Face: https://huggingface.co/YOON1v/Apex-2

(loads with transformers / vLLM via Qwen3MoeForCausalLM mapping)

Training

- Pretrain: 86.5B tokens (GH200 ×1 → ×2 with DiLoCo)

- SFT: ~2.5B tokens (code-heavy + math + instruction)

- DPO: tried it, scores dropped, so I dropped the checkpoint

Key numbers (SFT, greedy, chat template)

Benchmark

HumanEval 43.9

HumanEval+ 41.5

MBPP 56.3

MBPP+ 48.9

GSM8K (0-shot CoT) 32.4

MATH-500 21.0

IFEval (prompt strict) 44.7

MMLU (5-shot) 28.6

interesting comparison

With only ~0.087T pretrain tokens, the base model’s HumanEval+ matched Qwen2.5-1.5B (which used 18T).

Knowledge (MMLU) and math still lag far behind, as expected with the data gap.

What didn’t work

DPO (220k pairs, length-normalized) made answers much longer and hurt code / math / IFEval.

I stopped it and kept the SFT checkpoint. Full log is in the benchmark write-up.

Limitations (honest)

- English-centric (almost no multilingual ability)

- Weak knowledge → frequent hallucinations

- LiveCodeBench medium/hard is near zero

- 4k context only

Happy to answer questions about the MoE setup, DiLoCo, or why DPO backfired.

submitted by /u/Prestigious-Taste-63
[link] [留言]

来源:r/LocalLLaMA · reddit.com