DynamicTune 让 Qwen3.5-0.8B 在 ARC-Challenge 上以 42.15% 超过 2B 模型
A 0.8B model just beat a 2B model on ARC-Challenge (42.15%): Closed-form weight surgery beat multi-GPU SFT with 0 backprop (Independently verified on NVIDIA L4)
DynamicTune 通过闭式线性代数把大模型轨迹流迁移进小模型,在 ARC-Challenge 上让 Qwen3.5-0.8B 达到 42.15% acc_norm,超过同族 2B 基座的 41.10%。
A few days ago we shared the idea behind DynamicTune: transferring the trajectory flow from a larger teacher model directly into a smaller student via closed-form linear algebra in ~12 minutes on consumer hardware. Zero backpropagation, zero training tokens, zero gradient descent.
To eliminate local bias, we uploaded the unquantized FP16 checkpoint to Hugging Face, and TPN Bench (TaoFu Protocol) independently evaluated it on a datacenter NVIDIA L4 GPU using the official lm_eval 0.4.12 framework (coordinator run ce494664-d077-4ff1-8741-15cedabc434c). Huge thanks to TPN Bench for the cloud GPU compute!
Here are the independent numbers on full ARC-Challenge (1,172 items, zero-shot, greedy temp 0):
* Stock Qwen3.5-0.8B Base (unquantized BF16): 37.50% acc_norm (34.60% acc)
* 3-epoch SFT distillation (Mythos-0.8B, 25k Claude pairs, multi-GPU DDP): 38.10% acc_norm (35.80% acc)
* SFT + Model Soup Merge: 37.00% acc_norm (catastrophic forgetting)
* Stock Qwen3.5-2B Base (2.5x larger model, Q8): 41.10% acc_norm (37.80% acc)
* DynamicTune 0.8B Base (Ours, 4-anchor closed-form surgery): 42.15% acc_norm (40.19% acc)
WHY THIS IS COMPLETELY INSANE:
- A 0.8B model physically beat a 2.5x larger 2B model:
In LLM scaling, parameter count is supposed to be king. An 800M model is not supposed to beat an uncompressed 2B model on ARC-Challenge (42.15% vs 41.10%). By extracting trajectory dynamics from 4B and pulling them back into the student SwiGLU blocks, higher-order reasoning is compressed directly into edge weights.
- Zero backpropagation beat 25,000 SFT instruction pairs:
A recently published project (kmamine/merge-corrected-sft-distillation-Qwen-Mythos-0.8B) trained Qwen3.5-0.8B across 3 epochs on 25,000 Claude reasoning pairs on a multi-GPU cluster, reaching 38.10% before overfitting. DynamicTune reached 42.15% with zero gradient descent, zero loss functions, and zero training tokens.
- Ironclad 3.23-sigma statistical significance:
A delta of +4.65% across 1,172 questions with stderr +-1.44% gives a Z-score of 3.23sigma (p < 0.001). This is not prompt tuning noise or random variance.
- 12 minutes on consumer hardware vs datacenter verification:
The weight surgery was solved locally in ~12 minutes on an 8GB AMD RX 580 using layer-streaming (loading each layer in FP16, computing closed-form SVD deltas, and dumping to RAM). But the benchmark was conducted 100% in the cloud on datacenter NVIDIA L4 hardware via TPN Bench.
WHY PAST ATTEMPTS FAILED: THE SPECTRAL ENTROPY BARRIER
If you blindly apply weight deltas across all 24 layers of the student, the model collapses (+64.78% NLL explosion).
When we scanned all 24 layers calculating the normalized spectral entropy H (from 0.0 to 1.0) of the representation residuals:
* Layer 0 (H = 0.71): Clean semantic grounding. High receptivity to trajectory alignment.
* Layers 1-22 (H between 0.90 and 0.96): Chaotic superposition knots. In an 800M model with only 1024 dimensions, polysemantic features are crammed into dense superposition. Forcing linear updates here causes catastrophic interference.
* Layer 23 (H = 0.93): Pre-unembed boundary where features unpack toward vocabulary logits.
By restricting surgery to 4 sparse anchor blocks (layers 0, 7, 15, and 23) and using damped Levenberg-Marquardt Tikhonov pseudoinverse + adaptive spectral rank truncation, we protect the fragile superposition knots while imparting corrective trajectory velocity.
REPRODUCIBILITY & WEIGHTS
Everything is 100% open source and available to test right now:
* GitHub Repository: https://github.com/dsadawq3/DynamicTune
* Base Model (Safetensors): https://huggingface.co/F-Labs/Qwen3.5-0.8B-DynamicTune-Base
* GGUF Checkpoint (FP16): https://huggingface.co/F-Labs/Qwen3.5-0.8B-DynamicTune-Base-GGUF (Qwen3.5-0.8B-DynamicTune-Base-F16.gguf, SHA256: d77cf505108271d72f28298f20c2d158e7aeaf50cc22987db05a9a8973e08709)
To run inference locally with standard llama.cpp:
llama-cli -m Qwen3.5-0.8B-DynamicTune-Base-F16.gguf -p "Question: How does DNA replication initiate?\nAnswer:" -c 2048 -n 128
Special thanks to TPN Bench (TaoFu Protocol) for providing the independent datacenter NVIDIA L4 evaluation resources.
Clone the repo, run your own benchmarks, and test it yourself.
submitted by /u/AdventurousTwo6445
[link] [留言]
来源:r/LocalLLaMA · reddit.com