跳到正文
r/LocalLLaMA· /u/Henrie_the_dreamer·· 4 小时前AI 评分71

Cactus 发布 16.9MB 语音识别模型 Whistle

Whistle: speech to text in a 16.9MB file

AI 导读

Cactus Compute 发布 ASR 模型 Whistle,55m 参数(36m 激活)、CQ2bit 量化后仅 16.9MB,在 LibriSpeech test-clean 上 WER 为 4.31、test-other 为 10.49,对比 Whisper base 的 4.9 和 11.0,文件体积约为其九分之一、速度约六倍。

正文
Whistle: speech to text in a 16.9MB file

Hey all, we designed Cactus Whistle, an ASR model for ultra-small devices. It's not perfect, but mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish.

Remember, the goal at Cactus Compute isn't to achieve SOTA with scale, but to compress intelligence and bring them to smaller under-looked devices like budget phones, wearables, smart home and microcontrollers.

Whistle is 55m params (36m active) and CQ2bit quantised, amounting to a 16.9MB file that scores 4.31 WER on LibriSpeech test-clean and 10.49 on test-other, against 4.9 and 11.0 for Whisper base at 145.3MB. 21.4 on the FLEURS average against 24.5. SPGISpeech 7.65 and Earnings-22 19.01.

For the architecture, a log-mel front end and a convolution stem feed an audio encoder, and a Simple Attention + Hadamard MLP decoder reads it through gated cross attention at every layer. The decoder is laddered like Needle's, so every depth from 2 layers up is deployable.

Keyword biasing takes the names your users actually say and favours them during the beam search, which is what rescues a "Siobhan" or a "Krzysztof" from a model that was never told they exist. Word timestamps come from the decoder's own attention, so an app can highlight, seek or cut on a word.

Seventeen platforms are supported; macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32, Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly and a WASI component.

Please read more here: https://cactuscompute.com/blog/whistle

And let us know your thoughts!

submitted by /u/Henrie_the_dreamer
[link] [留言]

来源:r/LocalLLaMA · reddit.com