Transcription
Echophrase is model-agnostic: transcription backends are swappable, and the app ships with NVIDIA Parakeet-TDT 0.6B v3 as the default engine for fast, accurate speech-to-text. Qwen3-ASR 0.6B and Parakeet Lite 120M are also bundled.
How It Works
Echophrase captures audio from your microphone, processes it locally using the active transcription model, and types the transcribed text wherever your cursor is positioned.
Supported Languages
Language coverage depends on the active model:
- Qwen3-ASR 0.6B (international) — 52 languages, including Chinese
(Mandarin, Cantonese, and 20+ dialects), Japanese, Korean, Arabic, Hindi,
Thai, Vietnamese, Russian, and the major European languages. It also handles
mixed English/Chinese within a single sentence, and can convert its Chinese
output between Simplified and Traditional characters.
- Parakeet-TDT 0.6B v3 (default) — 25 European languages with automatic
detection.
- Parakeet Lite 120M — English only.
Dictating in Chinese, Japanese, Korean, or another Eastern language? Switch to
Qwen3-ASR 0.6B in Settings → Voice Model — the Parakeet family does not
cover these languages.
Accuracy
Transcription accuracy depends on:
- Model size: Larger models are more accurate but slower
- Audio quality: Clear audio produces better results
- Background noise: Quiet environments work best
See Model Selection to choose the right balance for your needs.