Supported AI Providers
Sokuji supports multiple AI providers for real-time speech translation. Each provider offers different capabilities, models, and pricing structures. Local (Offline) mode requires no API key or internet — it runs entirely on your device for free.
Setup Instructions
To use a cloud AI provider, obtain an API key from the provider's website and configure it in Sokuji's settings panel. Local (Offline) mode requires no API key — just download the models and start translating.
OpenAI
- GPT-4o Realtime Preview models
- 8 premium voice options
- Advanced turn detection modes
- Built-in noise reduction
- 60+ languages supported
- Template mode for custom prompts
Google Gemini
- Gemini 2.0 Flash Live models
- 30 unique voice personalities
- Automatic turn detection
- 35+ languages with regional variants
- Built-in transcription
- High token limits (8192)
PalabraAI
- Real-time WebRTC translation
- 60+ source languages
- 40+ target languages
- Low latency streaming
- Automatic audio processing
- Specialized for live translation
CometAPI
- OpenAI Realtime API compatibility
- Same voice and model options as OpenAI
- Alternative pricing structure
- Full feature parity
- Drop-in replacement for OpenAI
OpenAI Compatible API
- Any OpenAI Realtime API-compatible endpoint
- Custom base URL and API key
- Self-hosted and third-party services
- Flexible provider switching
- Private deployment support
Doubao AST 2.0
- End-to-end speech-to-speech translation
- Automatic voice cloning
- Low-latency simultaneous interpreting
- Supports zh, en, ja, id, es, pt, de, fr
- Auto and Push-to-Talk turn detection
- ByteDance cloud infrastructure
Soniox
- Real-time speech-to-speech translation
- Transcribe and translate 60+ languages in one API call
- Automatic language detection
- 12 multilingual voices
- Spoken translation or subtitles only
- Soniox cloud infrastructure
Local (Offline)
- Fully offline — no API key or internet needed
- 40+ ASR models covering 99+ languages
- 55+ translation language pairs + Qwen LLMs
- 136 TTS models across 53 languages
- CPU (WASM) and WebGPU acceleration
- Privacy-first — all data stays on-device
Choosing a Provider
OpenAI: Best for high-quality voice synthesis and advanced features
Gemini: Great for multilingual support and automatic processing
PalabraAI: Optimized for real-time translation with minimal latency
CometAPI: Cost-effective alternative to OpenAI with identical functionality
OpenAI Compatible: Connect to any OpenAI Realtime API-compatible service with a custom endpoint
Doubao AST 2.0: Doubao speech-to-speech simultaneous interpretation with automatic voice cloning
Soniox: Real-time speech-to-speech translation across 60+ languages with automatic language detection
Local Inference: Run everything on your device with no data leaving your machine — perfect for privacy-sensitive use cases
Need Help?
For setup guides, troubleshooting, and provider comparisons, visit our GitHub repository. GitHub Repository