C++ Apache-2.0

sherpa

Speech-to-text server framework with next-gen Kaldi

K

k2-fsa

Dernière activité 29 sept. 2026
k2-fsa/sherpa

997

étoiles

150

forks

110

issues ouvertes

asrcppctcend-to-end-asrpythonpytorchspeech-recognitiontransducerwebsocket

Ce README est souvent en anglais.

sherpa

sherpa is an open-source speech-text-text inference framework using PyTorch, focusing exclusively on end-to-end (E2E) models, namely transducer- and CTC-based models. It provides both C++ and Python APIs.

This project focuses on deployment, i.e., using pre-trained models to transcribe speech. If you are interested in how to train or fine-tune your own models, please refer to icefall.

We also have other similar projects that don't depend on PyTorch:

sherpa-onnx and sherpa-ncnn also support iOS, Android and embedded systems.

Installation and Usage

Please refer to the documentation at https://k2-fsa.github.io/sherpa/

Try it in your browser

Try sherpa from within your browser without installing anything: https://huggingface.co/spaces/k2-fsa/automatic-speech-recognition

Projets similaires

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

C++aarch64androidarm32
Kk2-fsa
15 k étoiles1,7 k

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

Pythonasrcode-switchconformer
PPaddlePaddle
12,7 k étoiles2 k

A PyTorch-based Speech Toolkit

Pythonasraudioaudio-processing
Sspeechbrain
11,8 k étoiles1,7 k