Skip to main content

Transcription

Echophrase is model-agnostic: transcription backends are swappable, and the app ships with NVIDIA Parakeet-TDT 0.6B v3 as the default engine for fast, accurate speech-to-text. Qwen3-ASR 0.6B and Parakeet Lite 120M are also bundled.

How It Works

Echophrase captures audio from your microphone, processes it locally using the active transcription model, and types the transcribed text wherever your cursor is positioned.

Supported Languages

Language coverage depends on the active model:
  • Qwen3-ASR 0.6B (international)52 languages, including Chinese (Mandarin, Cantonese, and 20+ dialects), Japanese, Korean, Arabic, Hindi, Thai, Vietnamese, Russian, and the major European languages. It also handles mixed English/Chinese within a single sentence, and can convert its Chinese output between Simplified and Traditional characters.
  • Parakeet-TDT 0.6B v3 (default) — 25 European languages with automatic detection.
  • Parakeet Lite 120M — English only.
Dictating in Chinese, Japanese, Korean, or another Eastern language? Switch to Qwen3-ASR 0.6B in Settings → Voice Model — the Parakeet family does not cover these languages.

Accuracy

Transcription accuracy depends on:
  • Model size: Larger models are more accurate but slower
  • Audio quality: Clear audio produces better results
  • Background noise: Quiet environments work best
See Model Selection to choose the right balance for your needs.