Python MIT

RealtimeTTS

Converts text to speech in realtime

K

KoljaB

Dernière activité 27 sept. 2026
KoljaB/RealtimeTTS

4 k

étoiles

409

forks

128

issues ouvertes

pythonrealtimespeech-synthesistext-to-speech

Ce README est souvent en anglais.

RealtimeTTS

PyPI Downloads GitHub release

RealtimeTTS is a Python text-to-speech library for applications that need to turn strings, generators, and LLM token streams into audio with low latency. It can play speech locally, stream chunks to another process, write WAV files, and fall back across multiple engines.

The project supports a broad engine matrix: local system voices, cloud APIs, free service wrappers, local neural models, and voice-cloning stacks.

Support RealtimeTTS

If RealtimeTTS saved you time, one GitHub star is a simple way to help make it more stable.

Stars improve visibility, and visibility brings more users, more real-world testing, more bug reports, more fixes, and better releases for everyone.

Demo

Short_RealtimeTTS_Demo.mov

For supported Windows and Linux systems with an NVIDIA GPU, QwenEngine is currently the recommended and preferred RealtimeTTS engine for high-quality, low-latency conversational speech. It offers multilingual Qwen3-TTS quality, x-vector and ICL voice cloning, native 24 kHz PCM streaming, fast cancellation, and the same engine either in-process or behind the production Qwen server.

In 10 warm Linux runs on our tuned RTX 4090 setup, the timeline was: about 35 ms engine TTFT, another 35 ms until RealtimeTTS emits its first PCM chunk, and about 10 ms of silence inside that chunk. Predicted audible onset was 80.9 ms and RTF was 0.108. These are orientation figures; measure the complete path on your target system.

RealtimeTTS 0.8.10 also provides QwenCpuEngine, using the maintained CPU-only native runtime with worker-pool improvements and streaming silence trimming. Choose one of these server installations in a fresh Python 3.11 or 3.12 virtual environment:

# NVIDIA GPU on Windows or Linux x86-64
python -m pip install "realtimetts[qwen-server]"
realtimetts-qwen-server --device gpu --clone-mode speaker_only --demo-voice
# CPU on Windows: example for 12 physical cores; adjust both worker counts
python -m pip install "realtimetts[qwen-cpu-server]"
realtimetts-qwen-server --device cpu --cpu-threads 6 --cpu-codec-threads 6 --cpu-core-split --cpu-stream-frames 1 --cpu-fused-attention --startup-buffer-ms 80 --onset-silence-profile qwen3_tts_12hz_0_6b_base_q8_v1 --demo-voice

Neither server extra needs Torch, a local CUDA Toolkit, or PortAudio. GPU mode still needs a compatible NVIDIA GPU/driver. CPU wheels support Windows x86-64, Linux x86-64 with glibc 2.35+, Intel macOS 13+, and Apple Silicon macOS 11+. x86-64 CPUs require AVX2/FMA/F16C/BMI2. CPU throughput depends on the machine; older x86 CPUs, Linux ARM CPU, Windows ARM64, and Metal are outside this release.

--device cpu --cpu-fused-attention --demo-voice alone keeps serial decoding. A positive --cpu-codec-threads enables overlap; --cpu-threads controls code generation and --cpu-stream-frames 1 selects 80 ms chunks. Six plus six is an example for 12 physical cores, not a universal default. On Windows, --cpu-core-split discovers the topology and separates the pools without changing process priority. It requires enough available cores in the highest performance class. Omit that Windows-only option on Linux/macOS. See the quick start for smaller CPUs and the full configuration used in the Windows latency checks. Studio starts with 0 ms additional browser buffering, adjustable under Sampling & options / Codec, transport & more options.

The optional --demo-voice prepares a public neutral cloning example, ready to select as demo-neutral. On Windows Ryzen 3900X, use --preset windows-3900x instead of --device cpu for the explicit overlap profile. Add --cpu-fused-attention to enable the optional CPU attention optimization. See the three-command CPU/GPU uv quick start.

Both servers offer browser playback at http://127.0.0.1:8080/studio, early em-dash speech, streaming segment controls, and language detection. See the QwenEngine guide for models, voices, authentication, language routing, and reproducing the deployed CPU settings.

CPU decoder overlap is available through explicit worker and chunk settings. See CPU scheduling and measured results. Windows Ryzen 9 3900X users can use the tested cmd launcher and fresh-venv guide.

Emotional Qwen demo from the video

The complete editable demo is in tests/faster_qwen_emotions.py, including the voice texts and playback loop. Its compact colored output is the default; add --verbose for native diagnostics. Warnings and errors remain visible. The packaged RealtimeTTS.qwen_emotions command is built from the same source. It keeps the original eleven emotional references/texts and the 0.6B Base Q8 speaker-only workflow. In an activated Python 3.11 or 3.12 environment:

git clone https://github.com/KoljaB/RealtimeTTS.git
cd RealtimeTTS
python -m pip install -e ".[qwen]"
cd tests
python faster_qwen_emotions.py

For CPU use .[qwen-cpu] when installing, then run python faster_qwen_emotions.py --device cpu. Local playback on Linux/macOS needs PortAudio (see below). The first run downloads the model pair if needed. See the emotional showcase guide for the complete virtual-environment setup, headless/server mode, and package-only commands.

InflectEngine is the documented lightweight alternative for one fixed English voice on CUDA or ONNX CPU.

Install

For the fastest local smoke test, install the system engine:

pip install "realtimetts[system]"

The system and other traditional engine extras use PyAudio. On Linux, install PortAudio headers before those extras:

sudo apt-get update
sudo apt-get install python3-dev portaudio19-dev

On macOS:

brew install portaudio

The local qwen, qwen-cpu, and Inflect extras use PyAudio/PortAudio for playback. Windows has prebuilt PyAudio wheels; on Linux and macOS install PortAudio first using the commands above. Use Python 3.11 or 3.12 for the Qwen workflows. Server-only installations do not need audio-device libraries.

Install realtimetts[qwen-server] (GPU) or realtimetts[qwen-cpu-server] (CPU) to expose the same native engine through an OpenAI-compatible HTTP API. The server provides /v1/audio/speech, dynamic voice registration, persistent voice latents, and watchdog-ready request/stall metrics on /health; it is headless and does not install PyAudio/PortAudio. The server defaults to loopback (127.0.0.1). LAN exposure requires a deliberate --allow-lan bind plus a built-in API key, or a trusted reverse proxy that terminates TLS and enforces authentication. CORS defaults to explicit localhost origins and rejects wildcard *; CORS is not an access control boundary. See the Qwen guide for deployment, protocol, licensing, and asset boundaries.

Sentence splitting defaults to stream2sentence's nltk+rule-based consensus mode. The normal install, including realtimetts[qwen], installs stream2sentence[nltk] but not Stanza or PyTorch. Add Stanza only when wanted:

pip install "realtimetts[stanza]"
# or
pip install "realtimetts[qwen,stanza]"

For cloud engines, local neural engines, CUDA, mpv, and current packaging caveats, see docs/installation.md.

First Audio

from RealtimeTTS import TextToAudioStream, SystemEngine


if __name__ == "__main__":
    stream = TextToAudioStream(SystemEngine())
    stream.feed("Hello from RealtimeTTS.")
    stream.play()

Use the if __name__ == "__main__": guard in scripts, especially on Windows and when using engines that start worker processes.

Streaming Text

feed() accepts an iterator, so text can arrive while audio is already playing:

from RealtimeTTS import TextToAudioStream, SystemEngine


def text_chunks():
    yield "This starts speaking quickly. "
    yield "More text can arrive while audio is already playing."


if __name__ == "__main__":
    stream = TextToAudioStream(SystemEngine())
    stream.feed(text_chunks())
    stream.play()

Use the same pattern with an LLM client by yielding only non-empty text chunks. See docs/llm-streaming.md.

Output

Write audio to a WAV file without local speaker playback:

from RealtimeTTS import TextToAudioStream, SystemEngine


if __name__ == "__main__":
    stream = TextToAudioStream(SystemEngine())
    stream.feed("Save this speech to a file.")
    stream.play(output_wavfile="speech.wav", muted=True)

For output devices, mpv playback, muted mode, callbacks, and chunk formats, see docs/output-and-files.md.

Features

  • Low-latency playback from strings, generators, and streamed model output.
  • Multiple engines with local, cloud, free-service, and neural model options.
  • Fallback engines for more resilient synthesis.
  • Sync and async playback with pause, resume, stop, and state inspection.
  • Text, audio, sentence, character, word-timing, and audio-chunk callbacks.
  • WAV output, muted synthesis, selected output devices, and volume control.
  • Voice switching and voice-cloning workflows where supported by the engine.

Engine Overview

Engine Type Install/status note Best first use
QwenEngine (recommended) Local native neural / HTTP server realtimetts[qwen] or realtimetts[qwen-server] with a matching native wheel High-quality multilingual realtime speech, voice cloning, and fast cancellation.
QwenCpuEngine Local CPU native / HTTP server realtimetts[qwen-cpu] or realtimetts[qwen-cpu-server] CPU-only Qwen on Windows/Linux x86-64 and Intel/Apple Silicon macOS; silence trimming and the same streaming controls.
InflectEngine Local lightweight realtimetts[inflect] Fast fixed English voice through PyTorch CUDA or ONNX CPU.
SystemEngine Local realtimetts[system] First local audio smoke test.
GTTSEngine Free service realtimetts[gtts] Simple network-backed speech.
EdgeEngine Free service realtimetts[edge], needs mpv Free streamed voices.
OpenAIEngine Cloud API realtimetts[openai] OpenAI TTS voices.
AzureEngine Cloud API realtimetts[azure] Azure voices and word timings.
ElevenlabsEngine Cloud API realtimetts[elevenlabs], needs mpv High-quality API voices.
CambEngine Cloud API realtimetts[camb] CAMB MARS API voices.
MiniMaxEngine Cloud API realtimetts[minimax] MiniMax cloud voices.
CartesiaEngine Cloud API realtimetts[cartesia] Cartesia API voices.
TypecastEngine Cloud API realtimetts[typecast] Typecast API voices.
ModelsLabEngine Cloud API realtimetts[modelslab] ModelsLab API voices.
CoquiEngine Local neural realtimetts[coqui] Local XTTS voice cloning.
PiperEngine Local executable realtimetts[piper], external Piper setup Fast local executable TTS.
StyleTTSEngine Local neural realtimetts[styletts], local checkout/assets StyleTTS experiments.
ParlerEngine Local neural realtimetts[parler] GPU local model experiments.
KokoroEngine Local neural realtimetts[kokoro] Local voices and timing support.
OrpheusEngine Local/API-style realtimetts[orpheus] Orpheus model workflows.
OmniVoiceEngine Local neural realtimetts[omnivoice] Multilingual voice cloning.
PocketTTSEngine / PocketTTSGpuEngine Local lightweight realtimetts[pockettts], realtimetts[pockettts-gpu] plus GPU fork CPU-oriented voice cloning, optional CUDA fork path.
NeuTTSEngine Local neural realtimetts[neutts], optional neutts-gguf Reference-audio voice cloning.
ZipVoiceEngine Local neural realtimetts[zipvoice], external checkout ZipVoice cloning/server demos.
LuxTTSEngine Local neural realtimetts[luxtts] LuxTTS voice cloning.
ChatterboxEngine Local neural realtimetts[chatterbox] Chatterbox prompt-audio voices.
SoproTTSEngine Local neural realtimetts[sopro] Sopro reference-audio voices.
SopranoEngine Local neural realtimetts[soprano] Soprano local synthesis.
MossTTSEngine Local neural realtimetts[moss], runtime assets MOSS-TTS experiments.
HiggsEngine Local HTTP PCM server realtimetts[higgs], separate SGLang-Omni server Stream Higgs Audio v3 PCM from a trusted server.

See docs/engine-selection.md before choosing an engine for an application. The engine-specific docs are being split out from the old README and source audit.

Documentation

  • Quick start: shortest working examples.
  • Installation: extras, platform setup, external tools, API keys, and known packaging mismatches.
  • Engine selection: engine matrix and selection guidance.
  • Feed and playback: feed(), play(), play_async(), pause, resume, stop, text state, and inline tags.
  • LLM streaming: provider-neutral streamed text patterns and latency tuning.
  • Output and files: WAV files, audio chunks, muted mode, output devices, mpv, buffering, and volume.
  • Forced alignment: optional word and character timing for engines without native timings.
  • Engine setup pages now link one focused page for each concrete engine source.
  • FAQ: legacy troubleshooting page while topic docs are being split out.

Legacy translated docs remain under docs/<locale>/ while English is refactored as the canonical source.

Server Example

The browser and WebSocket server example lives in example_fast_api/:

python -m pip install fastapi uvicorn websockets pyaudio
python example_fast_api/async_server.py

Open http://localhost:8000 or connect to ws://localhost:8000/ws.

RealtimeSTT is the speech-to-text counterpart for realtime voice input.

Contributing

Focused docs, tests, and engine fixes are easiest to review. During the docs refactor, keep English docs canonical and note mismatches between source, packaging, examples, and tests rather than hiding them.

License

RealtimeTTS source code is MIT licensed. Engine providers, model weights, voice data, datasets, generated audio, and third-party services can have separate terms. Read LICENSING_ADDENDUM.md and the relevant provider or model licenses before commercial use.

For the native Qwen/Inflect paths, qwentts.cpp and realtimetts-qwen-native are MIT-licensed; Qwen 0.6B Base/tokenizer and Inflect Micro-v2/ONNX are Apache-2.0. Model weights, voice latents, and reference audio are not bundled, and users are responsible for the rights to every voice or recording they provide.

Audio samples derived from the EARS dataset by Meta are licensed under CC BY-NC 4.0. See the original dataset terms for details.

Author

Kolja Beigel

Projets similaires

A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.

Pythonpythonrealtimespeech-to-text
KKoljaB
10,2 k étoiles861

Clone a voice in 5 seconds to generate arbitrary speech in real-time

Pythondeep-learningpythonpytorch
CCorentinJ
60,2 k étoiles9,4 k

Real-time text-to-speech with Qwen3-TTS

Python
Aandimarafioti
1,4 k étoiles196