Supported AI Providers

Sokuji supports multiple AI providers for real-time speech translation. Each provider offers different capabilities, models, and pricing structures. Local (Offline) mode requires no API key or internet — it runs entirely on your device for free.

Setup Instructions

To use a cloud AI provider, obtain an API key from the provider's website and configure it in Sokuji's settings panel. Local (Offline) mode requires no API key — just download the models and start translating.

OpenAI

Real-time Audio API
  • GPT-4o Realtime Preview models
  • 8 premium voice options
  • Advanced turn detection modes
  • Built-in noise reduction
  • 60+ languages supported
  • Template mode for custom prompts

Google Gemini

Gemini Live API
  • Gemini 2.0 Flash Live models
  • 30 unique voice personalities
  • Automatic turn detection
  • 35+ languages with regional variants
  • Built-in transcription
  • High token limits (8192)

PalabraAI

WebRTC Translation Service
  • Real-time WebRTC translation
  • 60+ source languages
  • 40+ target languages
  • Low latency streaming
  • Automatic audio processing
  • Specialized for live translation

CometAPI

OpenAI-Compatible API
OpenAI-compatible provider with same features
  • OpenAI Realtime API compatibility
  • Same voice and model options as OpenAI
  • Alternative pricing structure
  • Full feature parity
  • Drop-in replacement for OpenAI

OpenAI Compatible API

Custom Endpoint
Works with any OpenAI Realtime API-compatible endpoint
  • Any OpenAI Realtime API-compatible endpoint
  • Custom base URL and API key
  • Self-hosted and third-party services
  • Flexible provider switching
  • Private deployment support

Doubao AST 2.0

Speech-to-Speech Translation
  • End-to-end speech-to-speech translation
  • Automatic voice cloning
  • Low-latency simultaneous interpreting
  • Supports zh, en, ja, id, es, pt, de, fr
  • Auto and Push-to-Talk turn detection
  • ByteDance cloud infrastructure

Soniox

Speech-to-Speech Translation
  • Real-time speech-to-speech translation
  • Transcribe and translate 60+ languages in one API call
  • Automatic language detection
  • 12 multilingual voices
  • Spoken translation or subtitles only
  • Soniox cloud infrastructure

Local (Offline)

On-Device AI (Edge)
  • Fully offline — no API key or internet needed
  • 40+ ASR models covering 99+ languages
  • 55+ translation language pairs + Qwen LLMs
  • 136 TTS models across 53 languages
  • CPU (WASM) and WebGPU acceleration
  • Privacy-first — all data stays on-device

Choosing a Provider

OpenAI: Best for high-quality voice synthesis and advanced features

Gemini: Great for multilingual support and automatic processing

PalabraAI: Optimized for real-time translation with minimal latency

CometAPI: Cost-effective alternative to OpenAI with identical functionality

OpenAI Compatible: Connect to any OpenAI Realtime API-compatible service with a custom endpoint

Doubao AST 2.0: Doubao speech-to-speech simultaneous interpretation with automatic voice cloning

Soniox: Real-time speech-to-speech translation across 60+ languages with automatic language detection

Local Inference: Run everything on your device with no data leaving your machine — perfect for privacy-sensitive use cases

Need Help?

For setup guides, troubleshooting, and provider comparisons, visit our GitHub repository. GitHub Repository