Voice typing for Mac and Windows that turns speech into edited, formatted text in any app — plus screenshot questions and spoken edits behind one shortcut.
Nari Labs
Summary
Low-latency speech APIs from the team behind the open-source Dia model — text-to-speech and transcription priced well below the incumbents.
Screenshots
Description
Nari Labs is a Y Combinator-backed inference company building speech models and serving them behind an API tuned for latency rather than raw throughput. The team is known in the open-source community for Dia, a dialogue text-to-speech model with more than 2 million downloads, and for Narvatar, an avatar system; their public repositories have gathered over 20,000 GitHub stars.
The commercial products are two 1.7-billion-parameter models with a standard and a fast tier each:
Nari Qwen3-TTS 1.7B — text-to-speech at $5 per million characters, or $10 per million on the fast tier, which the company measures at under 50 ms server-side time-to-first-audio.
Nari Qwen3-ASR 1.7B — speech recognition at $0.06 per hour of audio, or $0.12 per hour on the fast tier, with under 40 ms server-side time-to-first-token.
Why it is interesting
The positioning is explicitly on the cost-latency frontier rather than on model size. Nari Labs reports placing on the Pareto frontier of the Coval voice benchmarks, claims first place on latency for both speech synthesis and recognition, and prices its text-to-speech at roughly a tenth of ElevenLabs. On density it reports running 80 concurrent live PersonaPlex-7B calls on a single NVIDIA H100 — the figure that matters if you are costing out a voice agent that must hold thousands of simultaneous conversations.
Beyond the shared API
Two services sit alongside the public endpoints. Dedicated inference provisions and tunes capacity for enterprise workloads that cannot share a queue, and Training fine-tunes the multimodal models on customer audio — relevant for domain vocabulary, accents and languages where general models degrade.
Nari Labs is aimed at developers building voice agents, telephony systems, dictation products and real-time assistants, where per-minute cost and the delay before the first syllable decide whether the conversation feels natural. Teams that want to self-host can start from the open models instead.
Reviews
Similar App Suggestions
Free AI audio and video transcription with timestamps, captions, and subtitle translation.
Voice typing that rewrites rough speech to fit the app you are in — Slack, Gmail, Cursor or a notes field — across 100+ languages.
Alibaba's AI music model: describe a mood, story or style in one line and get a finished song with lyrics, vocals and arrangement.
Speech-to-text service for meetings, interviews and video in 100+ languages, with an AI chat that answers questions against your transcripts.