ScenariosWhat are you trying to do?
These are presets, not modes you are locked into — every part of what they set can be changed afterwards.
What each one sets
Every one of them is the same three decisions, answered differently: whose audio gets translated, whether the result is spoken out loud, and what ends up on your screen.
understand-others
Understand what others say
Meetings, classes, talks, videos, streams — read a live translation of what you hear. Nothing is spoken back.
Who picks it: You are the listener, and the speaker is not going to meet you halfway.
what it sets
- whose audio
- the other side
- spoken aloud
- no
- on screen
- the translation of what they say
sokuji translatesbe-heard
Be understood in a meeting
They hear your translated voice through the virtual microphone, and carry on talking normally — no install, no plugin, no captions to read.
Who picks it: Anyone speaking in a language the room does not share, in a call they did not necessarily set up.
what it sets
- whose audio
- your microphone
- spoken aloud
- yes — a synthetic voice into the call
- on screen
- your words and the translation together
sokuji translatessubtitle-myself
Subtitle my own speech
Talks, streams, presentations — your audience reads translated subtitles while you speak normally. No synthetic voice is generated.
Who picks it: Presenters and streamers, where a second voice over your own would be worse than captions.
what it sets
- whose audio
- your microphone
- spoken aloud
- no
- on screen
- the translation, for your audience
sokuji translatestwo-way-voice
Two-way conversation
They hear your translated voice; you read subtitles built from what they say.
Who picks it: A real back-and-forth where neither side shares a language.
what it sets
- whose audio
- both sides
- spoken aloud
- yes — your side only
- on screen
- source and translation, both sides
sokuji translates
sokuji translatestwo-way-text
Two-way, subtitles only
Bilingual captions and meeting minutes — both sides as text, no synthetic voice at all.
Who picks it: Recorded meetings, note-taking, or any room where added audio would be unwelcome.
what it sets
- whose audio
- both sides
- spoken aloud
- no
- on screen
- source and translation, both sides
sokuji translates
sokuji translates“Any app” is literal. On the way out Sokuji is an ordinary microphone, so anything that lets you choose one can carry your translated voice — a browser tab, a recorder, streaming software, a game. On the way in it works from whatever your machine is playing, system audio included, so a video, a stream or a recording behaves exactly like a call.
Not every provider can do every scenario
It turns on one property: whether the provider can produce a voice, and whether it can be kept quiet. Some return text and nothing else, so they cannot speak your side into the call. Some always answer with audio, so they cannot be held to subtitles. Sokuji checks this when you choose the provider, so a scenario it cannot serve is closed off before the call rather than discovered during one.
available · not available with that provider
returns text
onlyalways
speakscan do
either
Understand what others say
Be understood in a meeting
Subtitle my own speech
Two-way conversation
Two-way, subtitles only
Understanding what others say is the one that works everywhere: that side is only ever turned into text, so whether the provider owns a voice never comes up.