Other's Audio

Translate the other side's audio in real time while you speak.

Overview

The Other's Audio feature enables real-time translation of what the other side is saying. While you speak and your voice is translated into Other's language, Sokuji can simultaneously capture the other side and translate it back into My language.

When to Use This Feature

  • Video conferences where participants speak different languages
  • Online meetings where you need to understand responses in real-time
  • International collaboration sessions
  • Language learning and practice sessions

How It Works

Sokuji uses a dual-client architecture to handle bidirectional translation without audio feedback loops:

Me channel (your voice)

Captures your microphone input and translates from My language into Other's language. The translated audio is played back to your meeting.

Other channel (the other side's voice)

Captures system or tab audio and translates from Other's language back into My language. The translation appears as text in the conversation view.

Platform Support

Other's Audio is available on all platforms supported by Sokuji:

PlatformSupportedCapture Method
Windows (Electron)✅System Audio Loopback
macOS (Electron)✅System Audio Loopback
Linux (Electron)✅PulseAudio/PipeWire Monitor
Browser Extension✅Chrome Tab Capture

How to Enable

Desktop App (Electron)

  1. Open the Settings panel
  2. Find the "Other's audio" section
  3. Toggle the switch to enable capture of the other side
  4. On Linux, select the audio output device to capture from
  5. Start your translation session
macOS Note: On macOS, you'll need to grant Screen Recording permission when first enabling this feature. Go to System Preferences > Security & Privacy > Privacy > Screen Recording.

Browser Extension

  1. Open the Sokuji extension panel
  2. Find the "Other's audio" section
  3. Toggle the switch to enable tab audio capture
  4. Select the output device for audio passthrough
  5. Start your translation session

Device Selection Guide

The device selection interface varies by platform. Here's what you'll see and what each option means:

Windows & macOS (Desktop App)

  • UI: Single "System Audio" option (no dropdown selection)
  • What it captures: All system audio output is captured automatically
  • Permissions:
    • Windows: No special permissions required
    • macOS: Screen Recording permission required (for audio loopback)
Note: Individual app audio selection is not available on these platforms.

Linux (Desktop App)

  • UI: Dropdown with multiple audio sinks + Refresh button
  • Options shown: Options show PulseAudio/PipeWire sink descriptions
  • Example options:
    • "Built-in Audio Analog Stereo" - Internal speakers
    • "HDMI Output" - External monitor audio
    • "USB Audio Device" - Connected USB headphones/speakers
  • What it captures: Audio output from the selected sink is captured
Tip: If your audio device isn't listed, click the Refresh button.

Browser Extension

  • UI: Dropdown with output devices + Refresh button
  • What the dropdown controls: The dropdown controls where passthrough audio plays (NOT what gets captured)
  • What gets captured: Current browser tab audio is captured automatically when enabled
  • Example options:
    • "Default" - System default output
    • "Built-in Speaker" - Laptop speakers
    • "USB Headphones" - Connected USB audio device
Important: Tab audio is captured automatically; the dropdown only selects where you hear the original audio (passthrough).

Platform Differences

The user experience varies slightly between platforms:

AspectDesktop AppBrowser Extension
Capture ScopeAll system audioCurrent tab only
Device SelectionLinux: Multiple sources available; Windows/macOS: Single "System Audio" optionOutput device for passthrough
Audio PassthroughNot needed (audio already playing)Automatic (tab is muted during capture)

Tips for Best Results

  • Use headphones to prevent audio feedback between your speakers and microphone
  • Ensure the other side's audio is clear and not distorted
  • The other side is translated to text only, to prevent echo
  • The Me and Other channels use the same AI provider
  • Close unnecessary audio applications to reduce background noise capture

Frequently Asked Questions

Why is the other side translated to text only?

To prevent audio feedback loops, the other side's translation is displayed as text rather than spoken aloud. This ensures you can read the translation without creating an echo in the meeting.

Can I capture audio from a specific application?

Currently, the desktop app captures all system audio. Application-specific audio routing is not supported. The browser extension only captures audio from the current tab.

Why do I need Screen Recording permission on macOS?

macOS requires Screen Recording permission for any application that captures system audio. This is a security feature of the operating system.

What happens if I use the same device for speaker output and Other's audio capture?

Sokuji will show a warning if you select conflicting devices, as this could create audio feedback. Choose different devices or use headphones to avoid this issue.