Open-source, fully local voice AI studio for macOS, Windows and Linux — voice cloning, voice design, dubbing, dictation and transcription in 646 languages.
Text-to-Speech
Turns written text into natural spoken audio, with control over voice, pace and delivery.
20 apps, 2 skills and 3 MCP servers tagged Text-to-Speech.
Apps
Free, fully on-device voice loop for AI coding agents — talk to Claude Code, Cursor or Codex and hear them answer back out loud.
Open-source inference engine that runs LLMs, vision and speech models fully on-device — no API keys, no cloud calls, from Swift to Godot.
A router for voice models: 56 speech and language models benchmarked across 10 languages, with each session sent to the one that actually wins.
Personal intelligence agent that tracks the topics you name across news, filings, forums and podcasts, folds duplicate coverage into one card, and briefs you by feed, push, email, RSS or audio.
Describe a video and get a complete edit — script, scenes, voiceover, and music — that you can refine by chatting.
Turn blog posts, scripts, and long recordings into narrated, captioned videos without editing software.
Listen to anything — documents, articles, PDFs, and email — in natural voices, at up to several times normal speed.
Enterprise speech infrastructure — fast, accurate transcription, text to speech, and a full voice agent API.
Voice AI that reads and expresses emotion — speech models trained on human vocal expression, not just phonemes.
Ultra-low-latency speech models built for real-time voice agents, where every millisecond is audible.
Studio-grade text to speech and voice cloning, with an open community library of thousands of voices.
Wondercraft makes studio-quality podcasts with AI.
Resemble AI provides high-quality voice cloning.
PlayHT offers ultra-realistic AI voices and cloning.
D-ID generates photoreal talking head videos.
Captions is an AI video creation and editing app.
Creates videos using AI avatars from text scripts
Build audio experiences at scale with Murf's AI voice generator
Generate lifelike spoken audio with AI tools
Skills
MiniMax's official skill for generating text, images, video, speech and music from the terminal through the mmx CLI, with agent-safe non-interactive flags.
NVIDIA's official skill for deploying and operating Nemotron Speech (Riva) NIMs — ASR, text-to-speech and translation, cloud-hosted or self-hosted on your own GPUs.
MCP servers
ElevenLabs' official MCP server — generate speech, clone voices, transcribe audio, and build voice agents.
Local speech MCP server bundled with VoiceStudio — generate speech, clone a voice from reference audio, and transcribe in 646 languages, all on your own machine with no API key.
MCP: Deepgram
by Deepgram
Deepgram's official MCP server, bringing speech-to-text, text-to-speech and audio intelligence into Claude Code, Cursor and Windsurf.
Related tags
Tags that appear alongside this one, ranked by how often.
