过度拟合推理引擎的兴起:Strata、ninfer、DwarfStar 等专用运行时涌现
The Rise of Overfit Inference Engines
一批极窄用途的推理运行时正在出现,包括 Strata、ninfer、DwarfStar、Splash、llamAmpere、gufo 等,它们放弃 llama.cpp/vLLM 擅长的通用性,围绕少量模型甚至单一硬件平台(如 Strix Halo)做优化。
| There seems to be a whole category of extremely narrow inference runtimes appearing: Strata, ninfer, DwarfStar, Splash, llamAmpere, gufo, etc. They deliberately give up the thing llama.cpp/vLLM are great at - generality - and optimize around a small number of models and It seems that general runtimes for compatibility, disposable overfit runtimes for maximum performance is going to be the norm forward. This is actually another good step in helping the democratization and decentralization of intelligence (models and runtimes both) and extracting more out of existing hardware where it doesn't have to be beautiful, well written, as long as it gets maximum output from one particular configuration. Curious if people think this the future/norm. submitted by /u/carteakey[link] [留言] |
来源:r/LocalLLaMA · reddit.com