modelscope/FunASR
#141 of 803 by Stars among GitHub repos
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
About
Industrial speech recognition toolkit for offline, streaming, and edge deployment.
No local setup? Open the Colab quickstart to transcribe a public sample or upload your own audio in a browser.
Flagship model — Fun-ASR-Nano (LLM-ASR for Chinese, English, and Japanese, plus Chinese dialect groups and regional accents; needs a GPU):
For the separate 31-language checkpoint, use Fun-ASR-MLT-Nano-2512. Language coverage is checkpoint-specific, so Nano and MLT-Nano should be treated as distinct model choices.
On CPU (or for five-language ASR plus emotion and audio-event tags), use SenseVoiceSmall. The pipeline below composes SenseVoiceSmall with FSMN-VAD and CAM++; diarization is provided by…
Excerpted from github.com/modelscope/FunASR
Latest metrics
| Stars | 20.3k | 2026-09-12 |
|---|---|---|
| Forks | 2.0k | 2026-09-12 |
| Commits | 5.9k | 2026-09-12 |
| Releases | 64 | 2026-09-12 |
| Watchers | 118 | 2026-09-12 |
| Open issues | 25 | 2026-09-12 |
| Open PRs | 3 | 2026-09-12 |