Speak your own language.
Be heard in theirs.

Your voice goes out in their language. What they say comes back as subtitles in yours. Sokuji is a virtual microphone between you and your meeting app, so the setup stays on your side — they just talk normally, with nothing to install and no link to join.

Free & open source · on-device AI · no subscription
Sokuji

Install the app, or add it to your browser

The desktop app is not tied to a list of supported websites. It registers as a microphone on your machine, so anything that can choose a microphone can use it.

Desktop app — system level

Sokuji appears as an ordinary microphone in your operating system. Meeting apps, browsers, recorders, streaming software, games — if it lets you pick a mic, it can hear your translated voice.

Coming the other way it asks permission of nobody: Sokuji reads whatever your machine is playing, system audio included, so a video, a stream or a recording is subtitled the same way a call is.

Linux / built in, on PulseAudio or PipeWire
macOS / driver ships with the app · Windows / via VB-CABLE

Browser extension

Nothing installed on the system, no audio routing to set up. Purpose-built for the meeting sites it supports.

Zoom · Google Meet · Microsoft Teams · Slack · Discord · Jitsi Meet · Whereby · Gather · Yandex Telemost

Inside those pages your translated voice is offered to the site as a microphone, and the tab's own audio is what it translates back. Anywhere else — another app, a video outside the browser, system audio — is the desktop app's job.

What are you trying to do?

Sokuji is built so the cost of crossing a language falls on the person who wants to be understood. That is what we built first. But you are not always that person — sometimes you are in a lecture, listening to a recording, or on a call somebody else set up.

Understand what others say

Meetings, classes, talks, videos, streams — read a live translation of what you hear.

them speaking → you read

Subtitle my own speech

Talks, streams, presentations — your audience reads translated subtitles, no synthetic voice.

you speaking → they read

Two-way conversation

They hear your translated voice; you read subtitles built from what they say.

both directions · voice out, text in

Two-way, subtitles only

Bilingual captions and meeting minutes — both sides as text, no synthetic voice at all.

both directions · text only

What it does during a call

Anyone can pipe speech through a model. Staying intelligible over a whole meeting — keeping pace, knowing when you have finished a thought, not echoing yourself — is audio engineering.

Speak in your own voice

Record a short clip once and your translated speech carries your voice instead of a stock one — or pick one from the provider’s library. Which voices you can choose from depends on the provider you are on.

Adaptive speech speed

The same sentence takes longer in some languages than others, so at a fixed rate the translation falls further behind with each one. Sokuji speeds the delivery up just enough to stay level.

Semantic turn detection

It waits for a finished thought, not merely a pause — so you are not cut off mid-sentence.

Your words, your terms

System instructions keep names, product terms and in-house jargon from being translated into mush.

Subtitle overlay

A separate window you can drag and resize, or captions laid over the meeting page itself.

Works with the network off

Sokuji detects your GPU, downloads the inference engine built for it, and runs recognition, translation and speech on your own machine — no key, no account, no meter. On-device is a first-class path here, not a fallback for when the connection drops.

Also: push-to-translate, noise suppression, echo cancellation, partial-transcript translation, audio buffer tuning, swapping audio devices mid-session, the virtual microphone itself, and exporting a conversation as .txt or .json.

What we carry for you — and what stays yours

Sokuji is free and open source, and stays that way. Separately, a Kizuna AI account can hold the key and settle the bill with a few AI vendors, for people who would rather not deal with developer consoles. We do not run the models — and it is a service you may never touch. Every convenience has an exit beside it.

we carry this

One account, no consoles

We issue the credential and settle the bill. OpenAI, Doubao or Soniox still runs the model.

One bill, metered

We handle the payment plumbing and charge you for what you actually used.

Cloud models when you need them

The fastest and most accurate translation available today, on demand.

A packaged app

Desktop builds and a browser extension, installed in a couple of clicks.

you can always do it yourself

Bring your own key

Ten-plus providers supported, each with a written walkthrough and screenshots.

Pay the provider direct

Leave us out of it entirely. Sokuji works without a Kizuna AI account at all.

Run it on your own machine

On-device inference, free, offline-capable, and shipped as a first choice.

Build it from source

Clone the app, read it, change it, ship your own build. AGPL-3.0.

35+
translation languages
30
interface languages
9
meeting sites in the extension
10+
AI providers, plus your own endpoint

Your audio does not route through us

Even when the key comes from your Kizuna AI account, audio goes from your device straight to the vendor running the model — and if the model you chose is the on-device one, it does not go anywhere at all. We issue the credential and meter the spend, and the only thing we store about you is an email address.

Every platform, and what each one asks of you

Windows

Code-signed, so it installs without the usual warnings — thanks to SignPath Foundation, who sign it for free. Routing uses VB-CABLE.

Download →

macOS

The audio driver ships with the installer. The app is unsigned, so Gatekeeper will stop you — a free project cannot carry Apple’s yearly fee.

Download →

Linux

The virtual microphone is created by Sokuji itself through PulseAudio or PipeWire — nothing else to install.

Download →

If you would donate an Apple Developer account for signing and notarisation, we would use it → support@kizuna.ai

Why it works this way

Most of the questions we get are really questions about the choices behind the product. Here they are, answered plainly.

Do you get my audio or my meeting content?

No. Your audio goes from your device straight to the AI provider you chose. This holds when the key comes from a Kizuna AI account too: we issue it and count usage, and the vendor does the rest. The only user data we hold is an email address, used to verify you and to grant trial credit.

Why should I believe that?

Partly, you shouldn’t — so here is the line. The app is AGPL-3.0 licensed, and the app is the piece that decides where your audio goes: all of it is at github.com/kizuna-ai-lab/sokuji, you can build it yourself, and you can watch it on the wire. Our account service is not open source, so what it does with an email address and a usage counter is a promise, not a proof. We would rather mark that line than blur it.

Why doesn’t the other person need to install anything?

Because we think crossing a language is the job of whoever wants to be understood, not of everyone listening. Sokuji translates on your side and speaks through a virtual microphone, so what reaches them is ordinary call audio.

Can I use it to translate YouTube in real time?

You can, and nothing stops you — Sokuji translates whatever audio reaches your machine, a video included. It is just not what we build for. A video watched by a hundred thousand people, translated live, is the same work done a hundred thousand times; translated once by the publisher, it is done once. Creators do use Sokuji to produce those translations.

Do I have to use your AI service?

No. Bring your own key for any supported vendor, point Sokuji at any OpenAI-compatible endpoint, or run on-device for free. The Kizuna AI account exists for people who would rather not deal with developer consoles — not to lock anyone in.

Is local AI just a fallback for when I’m offline?

No — it is a first-class path and has always been free. Models keep getting smaller, runtimes keep improving, and consumer accelerators keep getting faster. Nobody should have to hand their conversation to a large AI company to get AI at all.

What does Sokuji cost?

Sokuji is free and open source, and on-device inference is free with it. Only the optional Kizuna AI account costs anything, and it is metered by what you actually use — never a monthly subscription, and a quiet month costs nothing. Top-up balance is valid for a year: that bound is bookkeeping, not a nudge to spend. Bring your own key or run locally and you never pay us at all.

Why does macOS warn me, and Windows not?

Both come down to code signing, which costs money. Windows is signed thanks to SignPath Foundation, who sign free software at no charge. Apple has no such programme and charges a yearly developer fee, which a project with no revenue cannot carry — so the macOS build is unsigned, Gatekeeper says so, and the workaround is documented rather than hidden. Donations of an Apple Developer account are welcome.

Speak your own language.
Be heard in theirs.

free and open source · on-device ai included · no subscription