Turn a script into a natural, multi-speaker podcast or dialogue - with background music, stereo panning, and subtitles - in a few lines of Python.
- Multi-speaker dialogues with per-line left/right/both channel control.
- English and Spanish (plus more), thanks to pluggable TTS engines.
- Background music with automatic fade-in/out and ducking under speech.
- Voice cloning & emotion (via the Chatterbox engine) or unlimited random voices (via ChatTTS).
- Subtitles (
.srt/.vtt) generated automatically, perfectly timed to the audio. - WAV or MP3 output, plus a simple
podcast-ttscommand-line tool.
example-podcast_01.mp4
# 1. System audio tools (pick your OS)
brew install ffmpeg # macOS (Linux: apt-get install ffmpeg)
# 2. The library + the engine you want (see "Which engine?" below)
pip install "podcast_tts[chattts]" # English, unlimited random voices (default)
pip install "podcast_tts[chatterbox]" # English + Spanish, voice cloning, emotion
pip install "podcast_tts[kokoro]" # Fast & light, English + Spanish presets
pip install "podcast_tts[all]" # EverythingKokoro also needs
espeak-ng(brew install espeak-ng/apt-get install espeak-ng).
Pick one based on what matters most to you:
| Engine | Languages | Voices | Emotion | Speed | Best for |
|---|---|---|---|---|---|
| chattts (default) | English, Chinese | Unlimited random + your saved profiles | [laugh], breaks |
Medium | English podcasts, spinning up many distinct voices |
| chatterbox | 23 langs incl. Spanish | Clone any voice from ~10s audio | Yes (dial) | Slower (GPU recommended) | Spanish, cloning a real host, expressive delivery |
| kokoro | 8 langs incl. Spanish | 54 presets + blends | No | Fastest (CPU-friendly) | Quick, clean narration; low-resource machines |
You choose the engine when you create PodcastTTS(engine=...).
import asyncio
from podcast_tts import PodcastTTS
async def main():
tts = PodcastTTS(engine="chattts") # English default
await tts.generate_tts(
text="Hello! Welcome to our podcast.",
speaker="male1", # a premade voice
filename="hello.wav",
)
asyncio.run(main())import asyncio
from podcast_tts import PodcastTTS
async def main():
tts = PodcastTTS(engine="chattts")
dialogue = [
{"male1": ["Welcome to the show!", "left"]},
{"female2": ["Thanks for having me. [laugh]", "right"]},
{"male1": ["Today we talk about open source.", "left"]},
]
await tts.generate_dialog(dialogue, filename="dialogue.mp3", subtitles=True)
# -> dialogue.mp3 + dialogue.srt
asyncio.run(main())Use an engine that speaks Spanish (chatterbox or kokoro) and set the language.
import asyncio
from podcast_tts import PodcastTTS
async def main():
tts = PodcastTTS(engine="kokoro", language="es")
await tts.generate_tts(
text="Hola, bienvenidos al pódcast. Hoy hablamos de inteligencia artificial.",
speaker="ef_dora", # a Spanish Kokoro voice
filename="hola.wav",
)
asyncio.run(main())You can even mix languages in one dialogue by setting the language per line:
dialogue = [
{"Host": ["Welcome! Today we go bilingual."]},
{"Guest": ["Hola, gracias por la invitación.", "left", {"language": "es"}]},
]
await tts.generate_dialog(dialogue, filename="bilingual.mp3")music = [file_or_url, full_volume_seconds, fade_seconds, volume_under_speech]
await tts.generate_podcast(
texts=dialogue,
music=["intro.mp3", 10, 3, 0.3], # or a https:// URL (downloaded & cached)
filename="episode.mp3",
subtitles=True,
)The music plays at full volume, fades down under the dialogue, then fades back up and out.
Drop a clean 10-30s clip in your voices/ folder named after the speaker, or register it in code:
tts = PodcastTTS(engine="chatterbox", language="es")
tts.clone_voice("Ana", "samples/ana_reference.wav") # now "Ana" sounds like the clip
await tts.generate_tts(
"Hola, soy Ana y este es mi pódcast.",
speaker="Ana",
filename="ana.wav",
emotion=0.7, # 0.0 calm ... 1.0 dramatic
)podcast-tts say "Hello there" --speaker male1 -o hello.wav
podcast-tts dialog script.json -o show.mp3 --subtitles srt
podcast-tts dialog script.json -o show.mp3 --engine kokoro --language es \
--music intro.mp3 10 3 0.3script.json is just the dialogue list:
[
{"male1": ["Welcome to the show!", "both"]},
{"female2": ["Hola a todos.", "left", {"language": "es"}]}
]Prefer clicking to coding? Launch a small local web UI:
pip install "podcast_tts[demo,chattts]" # the demo + one engine
podcast-tts-demo # opens http://127.0.0.1:7860Two tabs: synthesize a single line, or paste a dialogue script and render a full podcast (with optional background music and a downloadable subtitle file).
-
ChatTTS ships three ready-to-use profiles:
male1,male2,female2. Any new name you use is generated once and saved tovoices/<name>.txtso it stays consistent. -
Chatterbox uses reference clips: put
voices/<name>.wav(or callclone_voice). Without a reference it uses its default voice. -
Kokoro uses preset ids (e.g.
af_heart,ef_dora,em_alex). Blend new ones:tts.engine.blend_voices("myvoice", {"ef_dora": 0.6, "em_alex": 0.4})
Each turn is a one-key dict: {"SpeakerName": [text, channel?, options?]}
text(str, required)channel(str, optional):"left","right", or"both"(default)options(dict, optional):{"language": "es", "emotion": 0.7}
tts = PodcastTTS(engine="chattts", language="en", speed=5, device=None)
await tts.generate_tts(text, speaker, filename="out.wav", channel="both",
language=None, emotion=None)
await tts.generate_dialog(texts, filename="dialog.wav", pause_duration=0.5,
normalize=True, subtitles=False, subtitle_format="srt",
language=None)
await tts.generate_podcast(texts, music, filename="podcast.wav", pause_duration=0.5,
normalize=True, subtitles=False, subtitle_format="srt",
language=None)The old API still works: from podcast_tts import PodcastTTS, plus generate_tts,
generate_dialog, and generate_podcast keep the same required arguments. New in 0.1.0:
the engine/language/emotion options, Spanish support, subtitles, and the CLI. The
default engine remains ChatTTS, so existing scripts behave as before.
pip install -e ".[dev]"
ruff check podcast_tts tests
pytest -qReleases are automated by .github/workflows/release.yml:
push a version tag and CI builds and publishes the package.
# 1. Bump the version in pyproject.toml (e.g. 0.1.0 -> 0.1.1)
# 2. Tag and push (the tag must match the pyproject version):
git tag v0.1.1
git push origin v0.1.1The workflow checks the tag matches the version, builds the sdist/wheel, and
uploads with skip-existing (so re-runs never clobber an existing release).
Authentication: the publish step uses a PYPI_API_TOKEN repository secret
(Settings → Secrets and variables → Actions). Create a PyPI API token scoped to
this project and store it there. To switch to
trusted publishing instead, drop
the password: line from the publish step, add permissions: id-token: write,
and register the publisher on PyPI.
Issues and pull requests are welcome on GitHub.
MIT - see LICENSE. Note the underlying engines have their own model licenses (ChatTTS, Chatterbox, Kokoro); review them for commercial use.
