Skip to content
Tech News
← Back to articles

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

read original more articles
Why This Matters

Nari Labs has achieved top rankings in voice AI benchmarks for both speech-to-text and text-to-speech, emphasizing low latency and high accuracy. This advancement is significant as it demonstrates more responsive and cost-effective voice AI models, which are crucial for improving user experience in voice-enabled applications.

Key Takeaways

TL;DR

Nari Labs leads Coval’s voice AI benchmark by sitting on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text. We also lead the latency-cost and quality-cost Pareto Frontier out of all publicly available models on the benchmark.

Coval is a leading provider of voice AI evaluation and benchmarks. They help speech AI agents perform better in production and publish one of the most widely cited benchmarks in the industry.

The Text-to-Speech (TTS) benchmark evaluates latency from text input to first audible chunk of audio (time-to-first-audio or TTFA) and Word Error Rate (WER). The Speech-to-Text (STT) benchmark evaluates latency from user’s finalize request to the final text output (time-to-final-segment or TTFS) and Word Error Rate (WER).

TTFA and TTFS are critical for voice agents, where latency can make a voice AI agent feel unresponsive. Low WER is an obvious key factor for model performance as well.

As of mid September 2026, Nari Labs tops both the Speech-to-Text and Text-to-Speech benchmarks. STT: #1 Latency, #2 WER. TTS: #2 Latency, #1 WER. Note that Coval’s benchmarks can fluctuate every 30 minutes*. We only include publicly available endpoints in our rankings and charts.

Speech-to-Text

Our Qwen3-ASR Fast model is ranked #1 in Time-to-Final-Segment (TTFS), at p50 of 44 ms and WER of 3.6%, placing #2 behind AssemblyAI’s Universal 3.5 Pro at 3.5%.

The pricing makes it even better. At $0.12 / hour, our Fast endpoint ties for the 2nd-lowest price among models with known public rates in Coval’s pricing directory. Universal 3.5 Pro costs 3.75× more, and Deepgram Nova 3 costs 2.4× more. Our Standard endpoint would be the cheapest at $0.06 / hour.

Text-to-Speech

... continue reading