zhifu gao
- modelscope/FunASR
- QwenAudio/SenseVoice
- modelscope/FunClip
- QwenAudio/Fun-ASR
- X-LANCE/SLAM-LLM
- modelscope/modelscope
- RVC-Boss/GPT-SoVITS
- DmitryRyumin/INTERSPEECH-2023-24-Papers
- 0xShug0/audio.cpp
- RapidAI/RapidASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
快速提取音视频内容,整理成一份结构化的markdown笔记
Fun-ASR speech recognition models, with native Hugging Face Transformers support for Fun-ASR-Nano and separate FunASR, vLLM and llama.cpp deployment paths.
📣 商用级开源语音自动识别程序库,开箱即用,全平台支持,中英文混合识别。A Cross-platform implementation of ASR inference. It's based on ONNXRuntime and FunASR. We provide a set of easier APIs to call ASR models.