跳到正文
r/LocalLLaMA· /u/Grand_Marionberry115·· 4 小时前AI 评分58

NihonSub:基于 Whisper-Large-v3 与 Groq/DeepSeek 的开源日语动漫实时字幕翻译引擎

I built an open-source real-time Japanese anime subtitle & translation engine powered by Whisper-Large-v3 + Groq / DeepSeek

AI 导读

作者发布开源工具 NihonSub,可将原始日语视频转成上下文双语字幕,并配有同步影院播放器。流水线用 ffmpeg 静音检测做 VAD 切分,Whisper Large-v3 在 Groq LPU 上转写。

正文
I built an open-source real-time Japanese anime subtitle & translation engine powered by Whisper-Large-v3 + Groq / DeepSeek

Hey r/LocalLLaMA,

Like many anime fans, I've always been frustrated by traditional MT engines (like Google Translate or base DeepL) when dealing with raw Japanese anime:

- They completely butcher Japanese honorifics, sentence-ending particles (-tteba, -zo, -desu wa), and character slang.

- They struggle with subject dropping (pro-drop grammar), translating pronouns inconsistently line-by-line.

- Cloud transcription APIs often choke on background music (OST), loud sound effects, and character screaming.

To solve this, I built NihonSub — an open-source tool and synchronized cinema player that turns raw Japanese video files into contextual bilingual subtitles.

🛠️ Architecture & Pipeline:

  1. Audio Extraction & VAD Chunking: Uses `ffmpeg` silence-detection to dynamically slice conversational utterances along natural speech pauses without chopping words in half.

  2. Speech-to-Text: Transcribes Japanese audio using OpenAI Whisper Large-v3 running on Groq LPUs for near-instant transcription speeds.

  3. Contextual LLM Translation: Feeds the transcript through DeepSeek / LLaMA-3 via Groq or OpenRouter with a specialized prompt that enforces anime nuance, honorific preservation, character tone, and simultaneous Hindi & English outputs.

  4. Synchronized Cinema UI: Custom WebVTT generator and video player with dual-subtitles, timestamp scrubbing, and full playback control.

💡 Why not just rely on standard NMT?

LLMs are far superior at resolving who is speaking to whom based on context and tone rather than naive literal dictionary lookup. With zero-cost free-tier APIs (Groq + OpenRouter free models), the entire pipeline runs without subscription costs.

Check out the demo video above!

- GitHub Repository: https://github.com/Abhishantpadam/NihonSub

- License: MIT

I'd love your thoughts on the pipeline, optimization ideas for local edge models (like running Whisper.cpp or local Ollama instances), or any feedback!

submitted by /u/Grand_Marionberry115
[link] [留言]

来源:r/LocalLLaMA · reddit.com