Free AI audio and video transcription with timestamps, captions, and subtitle translation.
VoiceStudio
Summary
Open-source, fully local voice AI studio for macOS, Windows and Linux — voice cloning, voice design, dubbing, dictation and transcription in 646 languages.
Screenshots
Description
VoiceStudio is a desktop application that puts an entire speech-AI pipeline on your own machine. Where hosted services bill per character and send your audio to their servers, VoiceStudio downloads the models and runs everything locally: no account, no API key, and no upload of the voices you clone.
One interface covers six jobs — cloning a voice from a short reference clip, designing a new voice from a description, dubbing video into another language, system-wide dictation behind a keyboard shortcut, transcription with speaker diarization, and long-form audiobook production. It claims coverage of 646 languages, and ships a batch queue plus vocal isolation for working through large jobs unattended.
Engines are swappable rather than fixed. Sixteen text-to-speech backends are bundled — OmniVoice is the default, alongside CosyVoice 3 and GPT-SoVITS — and eleven speech-recognition engines, defaulting to WhisperX with Faster-Whisper and MLX Whisper available. A built-in model catalogue handles downloading and switching between them.
VoiceStudio is also addressable by other software. The backend exposes an OpenAI-compatible REST API plus SSE and WebSocket transports, and mounts an MCP server at /mcp, so a coding agent such as Claude Code can generate speech, clone a voice or transcribe a file without leaving the terminal.
It is built with Tauri (Rust shell, React UI) and ships as a DMG, MSI and AppImage, with a Docker option. Requirements are macOS 13.3+ on Apple Silicon, Windows 10/11 x64, or Linux x86_64 with glibc 2.39+; 8 GB RAM minimum with 16 GB recommended, 10 GB of free disk, and an optional GPU (CUDA, Apple MPS/MLX, or ROCm). The application is licensed AGPL-3.0 — model weights keep their upstream terms, which is worth checking before commercial use. The project was previously named OmniVoice Studio.
Reviews
Similar App Suggestions
Hit record, and Voicenotes transcribes, summarises and indexes what you said — then answers questions about anything you have ever recorded.
A router for voice models: 56 speech and language models benchmarked across 10 languages, with each session sent to the one that actually wins.
Create stunning original music in seconds using AI. Make your own masterpieces, share with friends, and discover music from artists worldwide.
AI note-taker for real-world conversations — pairs recorder hardware and a desktop capture app with transcription in 112 languages and template-driven summaries.