Scenarios

What are you trying to do?

These are presets, not modes you are locked into — every part of what they set can be changed afterwards.

What each one sets

Every one of them is the same three decisions, answered differently: whose audio gets translated, whether the result is spoken out loud, and what ends up on your screen.

understand-others

Understand what others say

Meetings, classes, talks, videos, streams — read a live translation of what you hear. Nothing is spoken back.

Who picks it: You are the listener, and the speaker is not going to meet you halfway.

what it sets
whose audio
the other side
spoken aloud
no
on screen
the translation of what they say
be-heard

Be understood in a meeting

They hear your translated voice through the virtual microphone, and carry on talking normally — no install, no plugin, no captions to read.

Who picks it: Anyone speaking in a language the room does not share, in a call they did not necessarily set up.

what it sets
whose audio
your microphone
spoken aloud
yes — a synthetic voice into the call
on screen
your words and the translation together
subtitle-myself

Subtitle my own speech

Talks, streams, presentations — your audience reads translated subtitles while you speak normally. No synthetic voice is generated.

Who picks it: Presenters and streamers, where a second voice over your own would be worse than captions.

what it sets
whose audio
your microphone
spoken aloud
no
on screen
the translation, for your audience
two-way-voice

Two-way conversation

They hear your translated voice; you read subtitles built from what they say.

Who picks it: A real back-and-forth where neither side shares a language.

what it sets
whose audio
both sides
spoken aloud
yes — your side only
on screen
source and translation, both sides
two-way-text

Two-way, subtitles only

Bilingual captions and meeting minutes — both sides as text, no synthetic voice at all.

Who picks it: Recorded meetings, note-taking, or any room where added audio would be unwelcome.

what it sets
whose audio
both sides
spoken aloud
no
on screen
source and translation, both sides

“Any app” is literal. On the way out Sokuji is an ordinary microphone, so anything that lets you choose one can carry your translated voice — a browser tab, a recorder, streaming software, a game. On the way in it works from whatever your machine is playing, system audio included, so a video, a stream or a recording behaves exactly like a call.

Not every provider can do every scenario

It turns on one property: whether the provider can produce a voice, and whether it can be kept quiet. Some return text and nothing else, so they cannot speak your side into the call. Some always answer with audio, so they cannot be held to subtitles. Sokuji checks this when you choose the provider, so a scenario it cannot serve is closed off before the call rather than discovered during one.

available · not available with that provider

returns text onlyalways speakscan do either
Understand what others say
Be understood in a meeting
Subtitle my own speech
Two-way conversation
Two-way, subtitles only

Understanding what others say is the one that works everywhere: that side is only ever turned into text, so whether the provider owns a voice never comes up.

Speak your own language.
Be heard in theirs.

free and open source · on-device ai included · no subscription