Local (Offline) Setup Tutorial
Overview
Local (Offline) mode lets you run ASR (Automatic Speech Recognition), translation, and TTS (Text-to-Speech) entirely on your device. No API key, no internet connection, and no data leaves your machine. Everything is processed locally using WebAssembly and WebGPU.
No API Key or Internet Required
Privacy-first — all audio and text stay on your device. After downloading the models, Local (Offline) mode works completely offline. Your conversations never leave your machine.
1Select "Local (Offline)" as Provider
Open Sokuji and navigate to the Settings panel. Select "Local (Offline)" as your AI provider from the provider dropdown.
2Download an ASR Model
Choose an ASR model for your source language. Sokuji offers 42 offline/streaming models powered by sherpa-onnx WASM and WebGPU (Cohere Transcribe, Voxtral Mini, Whisper), covering 99+ languages. Models range from lightweight (~5MB) to high-accuracy (~2GB+). We recommend Cohere Transcribe or Voxtral Mini 4B Realtime for the best accuracy and speed.
3Download a Translation Model
Select a translation language pair. Choose TranslateGemma for the best quality and speed, Opus-MT for fast CPU-based translation, or Qwen LLMs via WebGPU for multilingual translation with higher quality.
4Download a TTS Model
Download a TTS model for spoken output. 136 models are available across 52 languages, powered by Piper, Coqui, Mimic3, and Matcha engines.
5Start Translating
Click the Start Session button. No internet connection is needed — translation happens entirely on your device in real time.
Additional Information
What is Local (Offline) Mode?
Local (Offline) is a fully on-device AI translation pipeline that runs ASR, translation, and TTS without any cloud services:
- Complete privacy — all audio and text processing stays on your device
- Works offline after initial model download
- No API keys, accounts, or subscriptions required
- Powered by WebAssembly (WASM) and WebGPU for native-like performance
- Free to use with no usage limits
Available Models
Local (Offline) mode supports a wide range of models for each stage of the translation pipeline:
- ASR: 42 models including Cohere Transcribe, Voxtral Mini 4B Realtime, Whisper, SenseVoice, Moonshine, NeMo, and Zipformer — covering 99+ languages
- Translation: 74 Opus-MT language pairs for fast CPU translation, plus TranslateGemma and Qwen LLMs via WebGPU for multilingual support
- TTS: 136 models across 52 languages using Piper, Coqui, Mimic3, and Matcha engines
Hardware Requirements
- CPU (WASM): Works on any modern browser — no special hardware needed
- WebGPU: Optional GPU acceleration for Cohere Transcribe, Voxtral Mini, Whisper ASR, TranslateGemma, and Qwen translation models — significantly faster with a dedicated GPU
- Recommended: 4GB+ RAM for comfortable model loading
- Storage: Models range from 5MB to 200MB+ depending on accuracy level
- Modern browser required: Chrome 113+, Edge 113+, or Firefox 120+ for WebGPU support
Model Catalog
ASR (Speech Recognition) Models
Choose a speech recognition model based on your source language, accuracy needs, and hardware capabilities.
Offline Models (CPU/WASM)
23 models that run on CPU via WebAssembly. No GPU required. Best for broad compatibility.
| Model | Languages | Size | Accuracy | Speed |
|---|---|---|---|---|
| SenseVoice (int8) | zh, en, ja, ko, yue | 238 MB | ★★★★☆ | ★★★★☆ |
| SenseVoice Nano (int8) | zh, en, ja, ko, yue | 265 MB | ★★★★★ | ★★★☆☆ |
| Moonshine Tiny EN (q) | en | 45 MB | ★★★☆☆ | ★★★★★ |
| Moonshine Tiny JA (q) | ja | 73 MB | ★★★☆☆ | ★★★★★ |
| Moonshine Tiny KO (q) | ko | 73 MB | ★★★☆☆ | ★★★★★ |
| Moonshine Base ZH (q) | zh | 142 MB | ★★★★☆ | ★★★★☆ |
| Moonshine Base JA (q) | ja | 142 MB | ★★★★☆ | ★★★★☆ |
| Moonshine Base ES (q) | es | 66 MB | ★★★★☆ | ★★★★☆ |
| Moonshine Base AR (q) | ar | 142 MB | ★★★★☆ | ★★★★☆ |
| Moonshine Base UK (q) | uk | 142 MB | ★★★★☆ | ★★★★☆ |
| Moonshine Base VI (q) | vi | 142 MB | ★★★★☆ | ★★★★☆ |
| NeMo Canary (int8) | en, es, de, fr | 208 MB | ★★★★☆ | ★★★☆☆ |
| NeMo FastConformer Multi | be, de, en, es, fr, hr, it, pl, ru, uk | 133 MB | ★★★★☆ | ★★★★☆ |
| NeMo FastConformer DE | de | 132 MB | ★★★★☆ | ★★★★☆ |
| NeMo FastConformer ES | es | 132 MB | ★★★★☆ | ★★★★☆ |
| NeMo FastConformer PT | pt | 132 MB | ★★★★☆ | ★★★★☆ |
| NeMo Parakeet TDT 0.6B (int8) | 25 EU languages | 671 MB | ★★★★★ | ★★★☆☆ |
| Dolphin Base CTC Multi | zh, ja, ko, th, vi, ar, hi, bn, ru | 105 MB | ★★★★☆ | ★★★★☆ |
| Whisper Tiny (ONNX) | 99+ languages | 154 MB | ★★★☆☆ | ★★☆☆☆ |
| WenetSpeech Yue U2++ | zh, yue, en | 135 MB | ★★★★☆ | ★★★★☆ |
| Omnilingual 300M v2 | 1147 languages | 367 MB | ★★★☆☆ | ★★★☆☆ |
| Zipformer RU (int8) | ru | 74 MB | ★★★★☆ | ★★★★☆ |
| Zipformer VI 30M (int8) | vi | 35 MB | ★★★☆☆ | ★★★★★ |
Streaming Models (CPU/WASM)
10 real-time streaming models. Transcribe as you speak with low latency.
| Model | Languages | Size | Accuracy | Speed |
|---|---|---|---|---|
| Zipformer EN Kroko | en | 71 MB | ★★★☆☆ | ★★★★★ |
| Zipformer FR Kroko | fr | 71 MB | ★★★☆☆ | ★★★★★ |
| Zipformer DE Kroko | de | 71 MB | ★★★☆☆ | ★★★★★ |
| Zipformer ES Kroko | es | 156 MB | ★★★☆☆ | ★★★★☆ |
| Zipformer ZH (int8) | zh | 77 MB | ★★★☆☆ | ★★★★★ |
| Zipformer ZH 2025 (int8) | zh | 167 MB | ★★★★☆ | ★★★★☆ |
| Zipformer RU Vosk (int8) | ru | 29 MB | ★★★☆☆ | ★★★★★ |
| Zipformer Multi (8 lang) | ar, en, id, ja, ru, th, vi, zh | 339 MB | ★★★☆☆ | ★★★☆☆ |
| Zipformer BN Vosk | bn | 94 MB | ★★★☆☆ | ★★★★☆ |
| NeMo CTC EN 80ms (int8) | en | 132 MB | ★★★★☆ | ★★★★★ |
WebGPU Models (GPU Required)
9 models accelerated by WebGPU including Cohere Transcribe and Voxtral Mini 4B Realtime (recommended) plus Whisper. Best accuracy and speed, requires a compatible GPU.
| Model | Languages | Size | Accuracy | Speed |
|---|---|---|---|---|
| Cohere Transcribe (q4) | 14 languages | ~2.1 GB | ★★★★★ | ★★★★☆ |
| Cohere Transcribe (q4f16) | 14 languages | ~1.5 GB | ★★★★★ | ★★★★☆ |
| Voxtral Mini 4B Realtime | 13 languages | ~2.3 GB | ★★★★★ | ★★★★★ |
| Whisper Tiny EN | en | 120 MB | ★★★☆☆ | ★★★★☆ |
| Whisper Tiny | 99+ languages | 122 MB | ★★★☆☆ | ★★★★☆ |
| Whisper Base | 99+ languages | 208 MB | ★★★★☆ | ★★★☆☆ |
| Whisper Small | 99+ languages | 586 MB | ★★★★☆ | ★★★☆☆ |
| Whisper Medium | 99+ languages | 680 MB | ★★★★★ | ★★☆☆☆ |
| Whisper Large V3 Turbo | 99+ languages | 759 MB | ★★★★★ | ★★☆☆☆ |
Translation Models
Choose TranslateGemma for the best quality and speed, Opus-MT for fast CPU translation, or Qwen LLMs for broad multilingual coverage.
TranslateGemma 4B (WebGPU) — Recommended
Google's purpose-built translation model with the highest quality and fastest speed among all local translation models. Supports 51 languages with any-to-any bidirectional translation. Requires WebGPU.
| Model | Languages | Size | Quality | Speed |
|---|---|---|---|---|
| TranslateGemma 4B (q4) | 51 languages | ~3.1 GB | ★★★★★ | ★★★★★ |
| TranslateGemma 4B (q4f16) | 51 languages | ~2.7 GB | ★★★★★ | ★★★★★ |
Arabic, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, English, Spanish, Estonian, Persian, Finnish, French, Gujarati, Hebrew, Hindi, Croatian, Hungarian, Indonesian, Icelandic, Italian, Japanese, Kannada, Korean, Lithuanian, Latvian, Malayalam, Marathi, Dutch, Norwegian, Punjabi, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Serbian, Swedish, Swahili, Tamil, Telugu, Thai, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Chinese, Zulu
Qwen LLM (WebGPU)
Large language models supporting 119–201+ languages. Higher quality but requires GPU. Best for flexible multilingual translation.
| Model | Languages | Size | Quality | Speed |
|---|---|---|---|---|
| Qwen 2.5 0.5B | 28 languages | ~400 MB | ★★★☆☆ | ★★★★☆ |
| Qwen 3 0.6B | 119+ languages | 919 MB | ★★★★☆ | ★★★☆☆ |
| Qwen 3.5 0.8B | 201+ languages | ~1.6 GB | ★★★★★ | ★★☆☆☆ |
| Qwen 3.5 2B | 201+ languages | ~1.8 GB | ★★★★★ | ★☆☆☆☆ |
Qwen 2.5: Japanese, Chinese, English, Korean, German, French, Spanish, Russian, Arabic, Portuguese, Thai, Vietnamese, Indonesian, Turkish, Dutch, Polish, Italian, Hindi, Swedish, Danish, Finnish, Hungarian, Romanian, Norwegian, Ukrainian, Czech, Estonian, Afrikaans
Qwen 3/3.5: 119–201+ languages (multilingual)
Opus-MT (CPU/WASM)
74 language pairs, ~110 MB each. Fast, lightweight, runs on any device. Best for single-direction translation.
| src ↓ tgt → | en | ja | zh | ko | de | fr | es | it | ru | ar | hi | vi | tr | uk | pt | nl | pl | da | sv | fi | cs | ro | hu | id | no | et | th | af | xh |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| en | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ||||||||
| ja | ✓ | ||||||||||||||||||||||||||||
| zh | ✓ | ||||||||||||||||||||||||||||
| ko | ✓ | ||||||||||||||||||||||||||||
| de | ✓ | ✓ | ✓ | ||||||||||||||||||||||||||
| fr | ✓ | ✓ | ✓ | ✓ | ✓ | ||||||||||||||||||||||||
| es | ✓ | ✓ | ✓ | ✓ | ✓ | ||||||||||||||||||||||||
| it | ✓ | ✓ | ✓ | ||||||||||||||||||||||||||
| ru | ✓ | ✓ | ✓ | ✓ | |||||||||||||||||||||||||
| ar | ✓ | ||||||||||||||||||||||||||||
| hi | ✓ | ||||||||||||||||||||||||||||
| vi | ✓ | ||||||||||||||||||||||||||||
| tr | ✓ | ||||||||||||||||||||||||||||
| uk | ✓ | ✓ | |||||||||||||||||||||||||||
| pt | |||||||||||||||||||||||||||||
| nl | ✓ | ✓ | |||||||||||||||||||||||||||
| pl | ✓ | ||||||||||||||||||||||||||||
| da | ✓ | ✓ | |||||||||||||||||||||||||||
| sv | ✓ | ||||||||||||||||||||||||||||
| fi | ✓ | ✓ | |||||||||||||||||||||||||||
| cs | ✓ | ||||||||||||||||||||||||||||
| ro | ✓ | ✓ | |||||||||||||||||||||||||||
| hu | ✓ | ||||||||||||||||||||||||||||
| id | ✓ | ||||||||||||||||||||||||||||
| no | ✓ | ||||||||||||||||||||||||||||
| et | ✓ | ||||||||||||||||||||||||||||
| th | ✓ | ||||||||||||||||||||||||||||
| af | ✓ | ||||||||||||||||||||||||||||
| xh | ✓ |
TTS (Text-to-Speech) Models
136 models across 52 languages. All run on CPU via WASM — no GPU required. Powered by Piper, Coqui, Mimic3, VITS, and Matcha engines.
TTS Settings
- Speech Speed — Controls how fast the synthesized speech is played. Adjust the speed slider to make the output faster or slower to suit your preference.
- Speaker ID — Controls the voice timbre of the synthesized speech. Many models support multiple speaker IDs, each with a different voice. Some multi-speaker models include voices of different genders, so you can choose a male or female voice by selecting different speaker IDs.
| Language | Models | Engine | Speakers | Size Range |
|---|---|---|---|---|
| af | 1 | mimic3 | 1 | ~94 MB |
| ar | 2 | piper | 1 | 81–95 MB |
| bg | 1 | coqui | 1 | ~71 MB |
| bn | 2 | mimic3, coqui | 1 | 71–94 MB |
| ca | 3 | piper | 1 | 39–95 MB |
| cs | 2 | piper | 1 | 81–95 MB |
| cy | 2 | piper | 1–7 | ~95 MB |
| da | 1 | piper | 1 | ~95 MB |
| de | 8 | piper | 1 | 39–81 MB |
| el | 2 | piper, mimic3 | 1 | 81–94 MB |
| en | 13 | piper | 1–904 | 81–97 MB |
| es | 6 | piper | 1–2 | 39–95 MB |
| et | 1 | coqui | 1 | ~71 MB |
| fa | 8 | piper, mimic3, matcha | 1 | 81–146 MB |
| fi | 2 | piper | 1 | 81–95 MB |
| fr | 8 | piper | 1–16 | 39–95 MB |
| ga | 1 | coqui | 1 | ~71 MB |
| gu | 1 | mimic3 | 1 | ~94 MB |
| hi | 3 | piper | 1 | ~95 MB |
| hr | 1 | coqui | 1 | ~71 MB |
| hu | 3 | piper | 1 | ~95 MB |
| id | 1 | piper | 1 | ~95 MB |
| is | 4 | piper | 1 | ~95 MB |
| it | 2 | piper | 1 | 39–95 MB |
| ka | 1 | piper | 1 | ~95 MB |
| kk | 3 | piper | 1+ | 39–81 MB |
| ko | 1 | mimic3 | 1 | ~94 MB |
| lb | 1 | piper | 1 | ~95 MB |
| lt | 1 | coqui | 1 | ~71 MB |
| lv | 1 | piper | 1 | ~95 MB |
| ml | 2 | piper | 1 | ~95 MB |
| mt | 1 | coqui | 1 | ~71 MB |
| ne | 3 | piper | 1+ | 39–95 MB |
| nl | 4 | piper | 1 | 39–95 MB |
| no | 1 | piper | 1 | ~95 MB |
| pl | 7 | piper | 1 | ~95 MB |
| pt | 4 | piper | 1 | 81–95 MB |
| ro | 1 | piper | 1 | ~95 MB |
| ru | 4 | piper | 1 | ~95 MB |
| sk | 1 | piper | 1 | ~95 MB |
| sl | 1 | piper | 1 | ~95 MB |
| sr | 1 | piper | multi | ~95 MB |
| sv | 2 | piper | 1+ | ~95 MB |
| sw | 1 | piper | 1 | ~95 MB |
| tn | 1 | mimic3 | multi | ~94 MB |
| tr | 3 | piper | 1 | ~95 MB |
| uk | 2 | piper | 1+ | 39–95 MB |
| vi | 4 | piper, mimic3 | 1–65 | 46–95 MB |
| yue | 1 | vits | 1 | ~114 MB |
| zh | 3 | vits, piper, matcha | 1–218 | 81–213 MB |
Troubleshooting
Common Issues
Model download fails: Check your internet connection for the initial download. Models are cached locally in IndexedDB, so subsequent uses work offline. Try clearing browser data if downloads are corrupted.
ASR not recognizing speech: Ensure you selected the correct ASR model for your source language. Try a larger model for better accuracy. Check that your microphone is working and permissions are granted.
Translation quality is low: Opus-MT models are optimized for speed over quality. For better results, try TranslateGemma or Qwen LLM models if your device supports WebGPU. Ensure you selected the correct language pair.
WebGPU not available: WebGPU requires a compatible browser and GPU. Fall back to WASM-based models (sherpa-onnx, Opus-MT) which work on any modern browser without GPU requirements.
Performance Tips
- Use WebGPU models when available for significantly faster processing
- Close other browser tabs to free up memory for model loading
- Choose smaller models for low-end devices or when speed is prioritized
- WASM models (sherpa-onnx, Opus-MT) are more compatible but slower than WebGPU alternatives
- Monitor browser memory usage — large models may require 2GB+ of browser memory
Need More Help? Visit our GitHub repository for community support, model compatibility guides, and the latest updates on Local (Offline) features.