用 FFT 替换 AdamW 优化器状态,显存减半微调 8B/70B 模型
We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?
有人尝试用 FFT 变换替代 AdamW 优化器状态来降低微调显存:将梯度转到频域、动态丢弃低影响频率后再逆变换,显存占用减少约 50%,代价是每步计算略慢。该方法面向消费级 GPU 上微调 8B 和 70B 模型时的 OOM 问题,团队正开放内部 Colab 环境和基线权重供他人验证。
Hey everyone,
Like most of you, we have been fighting constant OOM errors while trying to fine-tune 8B and 70B models on consumer GPUs. The AdamW optimizer states are always the biggest bottleneck.
We didn't want to rely on aggressive 8-bit quantization because we were seeing degradation in convergence, so we tried an experiment: tackling the optimizer states in the frequency domain.
The methodology:
Instead of storing the full gradients, we transform them using an FFT. This isolates the high-energy signal from the noise. We dynamically drop the low-impact frequencies and compress the state. When we inverse-transform back, it maintains the directional integrity but uses roughly 50% less VRAM.
The catch:
Running FFT operations adds compute overhead. It takes slightly longer per step, but the trade-off is completely avoiding OOM crashes and pushing batch sizes way up on standard hardware.
We are currently giving out access to our internal Colab environment and baseline weights to anyone who wants to poke holes in our math or try to break it.
We are really curious if anyone else here is exploring frequency-domain stuff or other non-quantization methods for VRAM reduction?
submitted by /u/Spectra-Global
[link] [留言]
来源:r/LocalLLaMA · reddit.com