Ship live translations
with confidence
A production-ready full-stack Node.js + React application for seamless EN↔RU↔UK live auto-detect translation with voice synthesis.
⚙
Installation
Set up the project locally with Docker, Redis, and LibreTranslate in minutes.
▦
Architecture
Understand the STT → Translation → TTS pipeline and real-time Socket.io communication.
▶
Live Translation
Stream from YouTube or microphone with automatic EN/RU/UK language detection and voice output.
📜
Biblical Simulator
Test the full pipeline with AI-generated biblical passages in King James, Church Slavonic, or Ukrainian style.
🎤
Voice Training
Clone custom voices from microphone recordings or YouTube videos using ElevenLabs IVC.
Prerequisites
●
Node.js 20+
Runtime for backend and build tools
●
Docker + Docker Compose
For Redis and LibreTranslate services
●
yt-dlp + ffmpeg
Required for YouTube audio extraction
●
ElevenLabs API Key
For speech-to-text and text-to-speech
Clone & Configure
git clone https://github.com/Pzharyuk/live-translator-node.git && cd live-translator-node
cp .env.example .env
Edit .env and set your API key:
ELEVENLABS_API_KEY=sk-your-key-here
ADMIN_PASSWORD=your-secure-password
Start Infrastructure
# Start Redis + LibreTranslate
docker compose -f docker-compose.local.yml up -d
# Wait for LibreTranslate to download language models (~500 MB)
docker logs -f $(docker ps -qf "name=libretranslate") 2>&1 | grep -i "running"
Start Backend
cd backend
npm install
npm run dev # nodemon watches for changes
Start Frontend
cd frontend
npm install
npm run dev # Vite hot-reload on localhost:5173
✓
You're all set!
Open http://localhost:5173 — log in with user / changeme and you will be redirected to /translate. Admin panel: http://localhost:5173/admin (admin password: admin123).
System Overview
Frontend
React 19 + Vite
Socket.io Client
Web Audio API
↔
Backend
Express + Socket.io
TypeScript
Service Layer
ElevenLabs
Scribe v2 (STT)
TTS Streaming
Voice Cloning
Translation
Google Translate (Cloud API)
LibreTranslate (self-hosted)
DeepL (premium API)
Claude / Anthropic (AI)
Redis
Feature Flags
Settings Store
Google Gemini
Biblical Simulator
Sermon Generation
Voice Training Text
DeepL
Free & Pro tiers
Auto endpoint detection
Church Directory
Central church registry
Live selector feed
Data Flow
1 Audio Input (Mic / YouTube / Simulator)
↓
2 PCM 16-bit LE @ 16kHz via Socket.io chunks
↓
3 ElevenLabs Scribe v2 WebSocket STT
↓
4 Commit Merge Buffer 2.5s VAD aggregation
↓
5 Translation Provider Google / LibreTranslate / DeepL / Claude
↓
6 ElevenLabs TTS Voice synthesis streaming
↓
7 Audio Playback Queued with 600ms pause
Key Architecture Decisions
Two-layer Language Detection
LibreTranslate's /detect endpoint returns 0-confidence for short Cyrillic phrases. The app uses script-based pre-detection (Unicode 0x0400–0x04FF = Cyrillic) combined with ElevenLabs Scribe's language_code output for reliable EN/RU/UK auto-detection.
VAD Commit Merging
Voice Activity Detection can fire aggressively on speaker breathing. Commits are buffered for 2.5 seconds before translation to merge fragments into meaningful phrases.
Feature Flag Merging
YAML config defaults are merged with Redis runtime overrides. Redis values take priority, falling back to YAML if Redis is unavailable.
API Key Hierarchy
Keys resolve in order: Runtime Cache → Redis → Config File → Empty. This allows hot-swapping keys without restarts.
Multi-church Isolation
Multi-church: Each church runs its own isolated deployment (own backend, Redis, database) at translate-<church>.hgministry.com. A church selector on the broadcast page lets listeners switch churches; the menu is fed by a central directory that updates live without touching running broadcasts. Audio, transcripts, and translations never cross between churches.
Connection Lifecycle
- Client sends
start_session with source type (mic or youtube) and optional voiceId
- Backend opens a WebSocket to
wss://api.elevenlabs.io/v1/speech-to-text/realtime
- For YouTube: spawns
yt-dlp | ffmpeg child processes to extract PCM audio
- For Microphone: awaits
audio_chunk events from the frontend
Audio Streaming
Audio chunks are sent to Scribe as JSON messages:
{
"message_type": "input_audio_chunk",
"audio_base_64": "UklGR..." // PCM 16-bit LE, 16kHz, mono
}
Scribe Responses
| Response Type | Meaning | Action |
partial_transcript |
Live partial text (speculative) |
Emitted as non-final transcript event |
committed_transcript |
VAD fired — complete phrase |
Buffered for commit merge window |
Commit Merge Buffer
After receiving a committed_transcript, the backend waits 2.5 seconds (COMMIT_MERGE_MS) to collect additional commits before translating. This prevents fragmented translations from aggressive VAD.
Stability Timeout
If VAD stalls (no new commits), a 3.5 second fallback timer (STABILITY_TIMEOUT_MS) fires to translate whatever new text has accumulated, preventing indefinite silence.
Text Validation
Before translation, text is validated against EN/RU/UK character regex patterns. This filters out hallucinated text from the STT model (common with silence or background noise).
Provider Chain
The system supports three translation providers with automatic fallback:
Default
LibreTranslate
Self-hosted, no API key required. Runs in Docker alongside the app. Best for privacy and cost.
Premium
DeepL
High-quality translations. Supports both free and paid API tiers. Auto-detects endpoint.
AI
Claude
Anthropic's Claude for context-aware translations. Uses claude-haiku-4-5 for speed.
Fallback Logic
1. Try primary provider (admin-selected)
2. If primary fails → try configured fallback
3. If fallback fails → try LibreTranslate (last resort)
4. If all fail → emit error event
Language Detection
The app uses a two-layer auto-detection approach:
Layer 1: Script-based Pre-detection
Before calling any translation API, the backend checks Unicode character scripts:
- Cyrillic characters (Unicode 0x0400–0x04FF) → if >50% of matched letters are Cyrillic, detected as Russian
- Latin characters → detected as English
- This avoids low-confidence results from LibreTranslate's
/detect endpoint on short text
Layer 2: STT Language Code
When the auto_language_detect flag is enabled, ElevenLabs Scribe returns a language_code with each transcript commit. The backend uses this to correctly route EN/RU/UK without relying solely on script detection.
Note: For LibreTranslate, both Russian and Ukrainian Cyrillic text is passed with source ru since LibreTranslate handles Ukrainian text acceptably via the Russian model. DeepL and Claude providers distinguish Ukrainian natively and handle uk as a proper source language.
Language Gating
Detected languages are checked against the admin-approved pool. If a detected language isn't in the allowed set, the translation is rejected to prevent hallucinated language outputs.
TTS Pipeline
After translation, the text is sent to ElevenLabs TTS:
const stream = await client.textToSpeech.stream(voiceId, {
text: translatedText,
model_id: "eleven_multilingual_v2",
output_format: "mp3_44100_128",
voice_settings: {
stability: 0.5,
similarity_boost: 0.75,
style: 0.0,
speed: 1.0,
use_speaker_boost: true
}
});
Audio Delivery
TTS audio is streamed to a Buffer, then emitted as a base64-encoded MP3 via the tts_audio Socket.io event.
Frontend Playback Queue
The frontend maintains an audio queue to prevent overlapping playback:
- Received
tts_audio events are queued
- Each segment plays to completion before the next starts
- A configurable pause (600ms default) is inserted between segments
- The pause duration is controlled by
tts_segment_pause_ms (adjustable in admin)
Microphone Input
- User selects "Mic" tab and chooses a TTS voice
- Browser captures audio via Web Audio API's
ScriptProcessor
- PCM 16-bit LE at 16kHz sample rate sent to backend via Socket.io
- Backend pipes audio to ElevenLabs Scribe v2 Realtime WebSocket
- Language auto-detected (EN/RU/UK), text translated and synthesized
- TTS audio returned and played back with inter-segment pauses
YouTube Input
- User pastes a YouTube URL (live stream or video)
- Backend spawns
yt-dlp | ffmpeg child processes
- Audio extracted as PCM stream (16kHz, 16-bit LE, mono)
- Piped to Scribe v2, same pipeline as microphone
- Stream ends when YouTube content ends or user stops
User Interface
The user view features a dark cavern theme with:
- Waveform visualizer — Canvas-based bar chart with orange gradient and cyan tips
- Transcript display — White translated text scrolls upward with fade masks
- Partial transcript — Shown in italic orange while STT is processing
- Source tabs — Toggle between Mic and YouTube (controlled by feature flags)
Congregation Viewer (/broadcast)
The public receiver page shows translated text only. Each line carries a small direction chip (RU → EN / EN → RU) so a viewer can see which way translation is currently running — it flips with the speaker — followed by the translation itself. The original-language line that used to sit beneath each translation was removed: it halved the space available for the text people actually read.
The translation is set in a responsive size — clamp(26px, 3.2vw, 34px) — so it stays comfortably readable on a phone held at arm's length and scales up on a laptop without turning a single sentence into a headline. Scrolling, the “Back to top” button, the fade masks, the live partial transcript, and the mute/TTS controls are unchanged.
How It Works
The backend uses yt-dlp and ffmpeg as child processes to extract audio from YouTube URLs:
yt-dlp (best audio) → ffmpeg (PCM 16kHz 16-bit LE mono) → Scribe v2
Supported Sources
- Live streams — Translates in real-time as the stream progresses
- Regular videos — Processes the full audio track
- Any URL supported by yt-dlp (YouTube, etc.)
Requirements
Both yt-dlp and ffmpeg must be installed and available in the system PATH. On macOS:
brew install yt-dlp ffmpeg
⚠
Feature Flag Required
YouTube input is controlled by the youtube_input feature flag. Enable it in the admin panel to show the YouTube tab in the user view.
Overview
The Biblical Transcript Simulator is an admin-only feature that generates biblical text passages using Google's Gemini API (gemini-2.5-flash), then routes them through the full translation pipeline. This provides a hands-free way to test STT → Translation → TTS without a live audio source.
Language Styles
| Language | Style | Example |
en |
King James English |
"In the beginning was the Word..." |
ru |
Church Slavonic Russian |
"В начале было Слово..." |
uk |
Traditional Ukrainian |
"На початку було Слово..." |
Flow
- Admin selects language (EN/RU/UK)
- Backend calls Gemini 2.5 Flash with streaming
- Gemini generates 6-8 biblical passages, 3-5 sentences each
- Stream is buffered until 140+ characters AND complete sentences
- Chunks emitted with 1800ms smooth pacing between them
- Each chunk flows through the standard pipeline:
- Emitted as
transcript (isFinal: true)
- Auto-translated via configured provider
- TTS synthesized and audio returned
- Frontend plays audio with standard inter-segment pause
💡
Feature Flag
Enable biblical_simulator in the admin feature flags panel. The Gemini API key is configured via the GEMINI_API_KEY environment variable or set at runtime in the admin API Keys panel.
Overview
Voice Training uses ElevenLabs' Instant Voice Cloning (IVC) API to create custom voices from audio samples. Once cloned, the voice appears in the voice selector immediately.
From Microphone
- Open the Voice Training section in the admin panel
- Click Generate Text to get an AI-generated reading passage (via Gemini) — gives the speaker natural, phonetically diverse text to read aloud
- Record multiple audio clips using your browser microphone while reading the generated text
- Provide a name for the voice
- Clips are uploaded to ElevenLabs IVC API
- Cloned voice is available for TTS immediately
- Click Preview Voice to hear the cloned voice speak a sample sentence via TTS
From YouTube
- Paste a YouTube URL in the Voice Training section
- Backend extracts N × 30-second clips via
yt-dlp + ffmpeg
- Clips are uploaded to ElevenLabs IVC API
- Resulting voice is stored in your ElevenLabs account
⚠
ElevenLabs Account
Cloned voices are stored in your ElevenLabs account, not locally. Ensure your plan supports voice cloning.
Concepts
| Concept | Description |
| Active Language Pair |
The current pair used for translation (e.g., EN ↔ RU, EN ↔ UK, or RU ↔ UK). Set by admin. |
| Available Languages |
The pool of languages viewers can select from (if user_language_selector is enabled). |
Admin Controls
- Change the active language pair via the admin panel
- Changes broadcast to all connected clients in real-time
- Manage the available languages pool for viewer selection
Viewer Selection
When the user_language_selector feature flag is enabled, viewers can override the admin-set language pair by selecting their own preferred languages from the available pool.
Overview
Two people can video call each other through the app, each speaking their own language. The app transcribes, translates, and synthesizes speech in real-time so each participant hears the other in their language.
Feature flag: Video call is gated behind the video_translation flag. Enable it in the admin panel or set video_translation: true in your YAML config.
How It Works
- Create a room — Person A selects their language, picks a TTS voice, and clicks "Create Room". A 6-character room code is generated.
- Share the code — Person A shares the room code with Person B (copy button provided).
- Join the room — Person B enters the code, selects their language and TTS voice, and clicks "Join".
- WebRTC connection — The app establishes a peer-to-peer video connection via WebRTC (signaled through Socket.io). Video flows directly between browsers.
- Audio translation — Each participant's microphone audio is simultaneously:
- Sent to the peer via WebRTC (but muted on their end)
- Captured as PCM chunks and sent to the backend via Socket.io for STT
- Translation pipeline — Each participant has their own independent Scribe STT session. Transcribed text is translated to the other participant's language, then synthesized via ElevenLabs TTS and sent back to the peer.
- Playback — The peer hears the TTS translation instead of the raw audio. Translated transcript is displayed below the video.
Architecture
Person A (Browser) Server Person B (Browser)
├─ getUserMedia ├─ Socket.io ├─ getUserMedia
├─ WebRTC P2P ═══video═══►│ (signaling) ◄═══ ├─ WebRTC P2P
│ │ │
├─ PCM chunks ──Socket.io─►├─ ScribeA(STT) │
│ │ ↓ translate │
│ │ ↓ TTS ───────────►├─ Plays TTS
│ │ │
│ Plays TTS ◄─────────────├─ ScribeB(STT) ◄───├─ PCM chunks
│ (remote video muted) │ ↓ translate │ (remote video muted)
└──────────────────────────┴────────────────────┘
Socket Events
| Event | Direction | Purpose |
video_create_room | C→S | Create a new room with language + voice |
video_room_created | S→C | Returns the 6-char room code |
video_join_room | C→S | Join an existing room |
video_room_joined | S→C | Sent to both participants, triggers WebRTC |
video_signal_offer/answer/ice | C↔S | WebRTC signaling relay |
video_audio_chunk | C→S | PCM audio for STT processing |
video_transcript | S→C | Transcript sent to the speaker |
video_translation | S→C | Translation sent to the listener |
video_tts_audio | S→C | TTS audio sent to the listener |
video_leave_room | C→S | Leave the room |
video_room_closed | S→C | Notify peer when other leaves |
Room Lifecycle
- Rooms are stored in Redis with key
video_room:{code} and a 4-hour TTL
- Maximum 2 participants per room
- When one participant disconnects, the other is notified and the call ends
- Scribe sessions are automatically cleaned up on disconnect
The Mac Audio Agent has moved to its own public repository:
github.com/Pzharyuk/live-translator-agent
It is a lightweight Node.js daemon that runs as a macOS LaunchAgent and streams microphone audio to the live-translator backend via Socket.io — eliminating the need to open a browser for the Remote Audio Source role.
Pre-shared key authentication
Any socket that emits register_audio_source must present the server's pre-shared key in the Socket.IO handshake (auth.agentPsk). This stops random clients from connecting to the backend and impersonating an agent.
- Server: set
AGENT_PSK (env var) — surfaces as auth.agent_psk in application.yaml. An empty value disables enforcement and logs a warning on every registration.
- Mac daemon: add
agentPsk to ~/.config/live-translator-agent/config.json (or set the AGENT_PSK env var — env wins).
- Browser
/audio-source: paste the key into the new Agent Pre-Shared Key field; it is stored in localStorage on that device only and travels in the handshake (never in event payloads).
- Mismatch behaviour: server logs
register_audio_source REJECTED ... invalid or missing PSK, emits agent_auth_error to the client, then disconnects.
Overview
/projector is a dedicated output route for the room screen — a projector, a confidence monitor, or a ProPresenter web element. It shows only the translated broadcast text: no original transcript, no live partial, no header, status pill, mute button, or schedule chrome. The newest line appears at the bottom and the whole stack rises as more text arrives.
The page is silent. It never plays TTS audio — sound comes from the room PA. This lets it run fullscreen on a booth machine indefinitely without fighting the sound system for control of playback.
The route itself requires no login, matching /broadcast, so it can be opened directly on a projector computer that has no admin session. Visibility is instead controlled by the projector feature flag.
⚠
Feature Flag Required
/projector is gated by the projector feature flag, which ships false in production. Enable it per-deployment in the admin Feature Flags panel before pointing a projector at the URL.
The admin Projector Output section is gated by the same flag: with the flag off the section is not rendered at all, and it appears the moment the flag is toggled on — no page reload needed.
Adaptive-Creep Scrolling
Rather than jumping line-by-line or scrolling at a fixed speed, the stack drifts continuously so the congregation reads a flowing column instead of a cursor jump:
- Base creep — a slow continuous drift,
18 px/s by default (?speed=).
- Catch-up — speed increases proportionally to how far behind the speaker the text has fallen, targeting a close-the-gap time of
tau seconds (?tau=, default 4, range 0.5–30). Also a live slider in the admin panel — see Catch-up below.
- Blur cap — speed never exceeds
130 px/s, so catching up never makes the text unreadable.
- Idle freeze — the moment nobody is speaking and there is no backlog left, motion stops completely rather than idling forward.
- Desync snap — a reconnect or a history replay after rejoin can leave the text far behind the live edge; once the backlog is large enough that crawling back would take minutes, the view snaps straight to the live edge instead.
- Centred while short — before the text fills the frame (the start of a service, or just after a clear) the stack is centred vertically rather than clinging to the top. Text is centred horizontally in both layouts.
Prayer & Song Pauses
When the broadcast is paused for prayer or song, the text fades out and a quiet message (“Prayer in progress” / “Song in progress”) fades in over the same spot. The scroll position is frozen, not reset, so when the pause ends the text fades back in exactly where the reader left off.
URL Parameters
All parameters are optional and independently clampable — an out-of-range value (typed by hand, or a leftover from a previous config) is clamped to the nearest valid bound rather than breaking the page.
| Parameter | Default | Description |
bg |
dark |
Backdrop: dark (app theme), transparent, or green (chroma-key). With layout=lower3, green fills only the band — everything above it stays transparent, so ProPresenter keys a defined lower-third region instead of the whole frame. |
layout |
full |
full fills the frame; lower3 confines text to a band pinned to the bottom — one line tall by default, with the text creeping continuously through it. |
size |
1 |
Font-size multiplier applied to the responsive base text size. |
band |
auto |
Lower-third band height. Left off (or auto) the band is exactly one line tall, derived from the same expression as the font size so it tracks size and the screen's aspect ratio. Give it a number to override in vh. Ignored when layout=full. |
speed |
18 |
Base scroll creep, px/sec. Clamped between 2 and 130. |
tau |
4 |
Seconds the engine targets to close any backlog before it falls back to the base creep. |
ProPresenter Setup
- Add the
/projector URL as a web element in ProPresenter, sized 1920×1080.
- Use
layout=lower3 so the text stays confined to the bottom band — one line tall unless band says otherwise — and bg=green to fill that band with a solid chroma-key green. The area above the band stays transparent, so only the band needs keying.
- Chroma-key the green in ProPresenter's compositing so only the text remains over your lyric/camera background.
⚠
Why green, not transparent
ProPresenter 7's web element has historically been unreliable about honouring page transparency — several versions composite it onto opaque black instead. bg=green is the recommended path for that reason; bg=transparent is available but should be verified on the actual playout machine before relying on it.
Admin “Projector Output” Panel
The admin panel's Projector Output section builds the launch URL and controls a running projector without editing query strings by hand:
- URL builder — layout, backdrop, text size, and band height (when in lower-third; leave the field blank for a one-line band), reflected live in the copyable link.
- Live Speed drag bar — a 2–130 px/s slider that retunes an already-open projector window instantly, so an operator can tune the creep by eye while watching the actual screen.
- Catch-up drag bar — the same live control for
tau, 0.5–30 s, labelled lower = text arrives sooner. On the one-line lower third this is the dominant remaining delay between the preacher speaking and the line being readable: the scroll aims to close whatever backlog exists within this many seconds, so 4 s means a newly arrived line takes about four seconds to rise into view and 1 s brings it up almost immediately. Pushed over the same BroadcastChannel as Speed, on its own message type, so the two sliders can never be read for each other.
- Copy link — copies the absolute
/projector URL with the current options encoded.
- Launch button — opens the projector in a new window. Once displays have been detected it launches fullscreen directly onto the chosen one instead.
- Close projector window — stops the output without hunting the fullscreen window down on the other screen. The request travels over
BroadcastChannel, not the handle window.open returned, so it still works after the admin page has been reloaded. There is no delivery receipt, so the button confirms only “Close sent” — it is harmless to press when nothing is open, and the panel never claims to know whether a projector is running.
- Detect displays — lists the physical displays attached to the current machine so you can pick one. Chrome only shows its permission prompt in response to a click, so this is a button rather than something that happens on page load: press it and choose Allow. Once granted, the list appears automatically on every later visit from that browser.
- Settings are remembered — layout, backdrop, text size, band height, speed, and catch-up are saved in the browser's
localStorage and restored on the next visit, so the projector only has to be configured once per booth machine. Saved values are re-validated on read and clamped to the same ranges the URL accepts. The selected display is deliberately not saved — screens come and go, so detection re-runs each time.
💡
If no display list appears
The panel always states which case you are in, directly under the “Displays attached to this machine” heading:
- “Click Detect displays…” — permission has not been asked for yet. Press the button and choose Allow.
- “Detection was dismissed…” — the prompt was closed without granting. Press the button again.
- “Display selection is blocked…” — permission was denied. Chrome will not ask again from a click; re-enable Window management in the site settings behind the icon at the left of the address bar.
- “Only one display is attached…” — detection worked and there is genuinely nothing to choose between. Connect the projector, then press Detect displays again.
- “Display selection needs a Chromium browser…” — Safari and Firefox do not implement the Window Management API. Copy the URL or use Open in new window and press F11 on the projector display.
⚠
Three limitations to know before service
A hand-opened window cannot be closed remotely. Browsers only let a script close a window a script opened — so Close projector window genuinely closes a projector started with Launch, but cannot close one opened by pasting the URL into a tab. That window is not left unchanged: it stops showing the service text, leaves fullscreen, and displays “Projector output stopped — you can close this window”, then drops out of the broadcast viewer count. Reload the page to resume output. Like the Live Speed bar, this only reaches projector windows in the same browser on the same machine.
Display list is local. The picker only enumerates displays physically attached to the machine the admin page is open on — it cannot see or drive a projector connected to a different computer. It also requires the browser's Window management permission, granted through the Detect displays button.
Live Speed and Catch-up bars are same-browser only. They push updates over BroadcastChannel, which only reaches a projector window opened from that same browser. A projector running on another machine keeps whatever values were in its ?speed= and ?tau= parameters at page load — the live bars cannot retune it.
Feature Flags
Feature flags control which features are visible and active in the application. Defaults are defined in config/application.yaml under the feature_flags section. Flags can be toggled at runtime via the Admin UI or the POST /admin/flags/:flag endpoint, and the changes are persisted in Redis, overriding the YAML defaults for all connected clients.
| Flag |
Default |
Description |
youtube_input |
true |
Enable YouTube live stream input as a broadcast source. |
mic_input |
true |
Enable microphone input from the admin browser. |
auto_language_detect |
true |
Automatically detect source language before translation. |
user_language_selector |
false |
Allow viewers to select their preferred language pair. |
audio_device_selector |
true |
Show audio device selection in the admin panel. |
video_translation |
true |
Enable real-time translation for video calls. |
video_voice_cloning |
false |
Premium feature: show Clone Voice button in video call lobby. |
remote_audio_source |
false |
Enable the /audio-source route for headless remote audio relay agents. |
agent_audio_source |
false |
Show connected remote agent audio sources in the admin panel. |
stream_input |
false |
Enable HTTP/Icecast audio stream input (admin can save a stream URL). |
auto_pause_songs |
false |
Automatically pause translation during detected worship songs. |
broadcast |
false |
Enable the /broadcast route for public receiver page. |
translate |
false |
Enable the /translate route for live translator mode. |
projector |
false |
Enable the /projector route for congregation screen & ProPresenter output. |
Programmatic API
Get all flags (merged defaults + Redis overrides):
GET /admin/flags
Response: { "flags": { "youtube_input": true, "auto_pause_songs": false, … } }
Get a single flag:
GET /admin/flags/:flag
Response: { "flag": "auto_pause_songs", "value": false }
Set a flag:
POST /admin/flags/:flag
Content-Type: application/json
{ "value": true }
Response: { "flag": "auto_pause_songs", "value": true }
Changes are persisted to Redis and broadcast to all connected Socket.IO clients via the feature_flags event, so the UI updates in real-time across all viewers and admins. When a flag is changed, all connected clients receive:
Socket.IO Event: 'feature_flags'
{ "youtube_input": true, "auto_pause_songs": true, … }
Special behavior: When auto_pause_songs is toggled to false at runtime, the releaseAutoPause() function is automatically called to stop the song detector and lift any active auto-pause immediately, preventing the broadcast from remaining stuck in a paused state.
File Structure
| File | Purpose |
config/application.yaml |
Base defaults for all environments |
config/application-local.yaml |
Local development overrides (localhost URLs) |
config/application-prod.yaml |
Production overrides (Docker service names) |
The APP_ENV environment variable (local or prod) determines which overlay file is loaded on top of the base config.
Full Configuration Reference
server:
port: 3001
cors_origin: "http://localhost:5173"
elevenlabs:
api_key: "${ELEVENLABS_API_KEY}"
default_voice_id: "kxj9qk6u5PfI0ITgJwO0"
tts_model: "eleven_multilingual_v2"
tts_settings:
stability: 0.5
similarity_boost: 0.75
style: 0.0
speed: 1.0
use_speaker_boost: true
stt_model: "scribe_v2"
anthropic:
api_key: "${ANTHROPIC_API_KEY}"
deepl:
api_key: "${DEEPL_API_KEY}"
libretranslate:
url: "http://libretranslate:5000"
api_key: ""
redis:
host: "redis"
port: 6379
password: ""
feature_flags:
youtube_input: true
mic_input: true
auto_language_detect: true
user_language_selector: false
audio_device_selector: true
video_translation: false
video_voice_cloning: false
broadcast: false
audio:
sample_rate: 16000
channels: 1
chunk_duration_ms: 250
translation:
source_lang: "auto"
target_lang_en: "en"
target_lang_ru: "ru"
provider: "libretranslate"
fallback: "libretranslate"
Environment Variable Interpolation
YAML values using ${VAR_NAME} syntax are automatically replaced with the corresponding environment variable at startup.
TTS Settings
Configure text-to-speech parameters for the ElevenLabs integration. Settings can be modified at runtime via the admin API and are persisted to Redis.
API Endpoints
GET /admin/tts-settings
Returns current TTS settings.
Response:
{
"settings": {
"stability": 0.5,
"similarity_boost": 0.75,
"style": 0.0,
"speed": 1.0,
"use_speaker_boost": true
}
}
POST /admin/tts-settings
Update one or more TTS settings.
Request Body:
{
"stability": 0.5,
"similarity_boost": 0.75,
"style": 0.0,
"speed": 1.0,
"use_speaker_boost": true
}
Response:
{
"settings": { ... updated settings ... }
}
Settings Reference
| Setting |
Range |
Default |
Description |
stability |
0.0 → 1.0 |
0.5 |
Voice stability: lower = more variable, higher = more consistent. |
similarity_boost |
0.0 → 1.0 |
0.75 |
How closely the output matches the voice sample; higher = closer match. |
style |
0.0 → 1.0 |
0.0 |
Exaggeration level of voice style (only on some voices); higher = more dramatic. |
speed |
0.5 → 2.0 |
1.0 |
Playback speed multiplier; 1.0 = normal, <1.0 = slower, >1.0 = faster. |
use_speaker_boost |
true | false |
true |
Enable speaker boost for enhanced clarity and volume consistency. |
STT Timing Settings
Control speech-to-text recognition timing and buffering behavior. Affects when transcripts are dispatched for translation.
GET /admin/stt-timing
Returns current STT timing configuration.
Response:
{
"settings": {
"commit_merge_ms": 2500,
"stability_timeout_ms": 2000,
"tts_segment_pause_ms": 0,
"max_accumulation_ms": 8000,
"vad_threshold": 0.5,
"vad_silence_threshold_secs": 1.5,
"min_speech_duration_ms": 100,
"min_silence_duration_ms": 100,
"flush_on_sentence_boundary": true,
"min_chars_before_dispatch": 40
}
}
POST /admin/stt-timing
Update one or more STT timing settings.
Request Body:
{
"commit_merge_ms": 2500,
"stability_timeout_ms": 2000,
"max_accumulation_ms": 8000,
...
}
Response:
{
"settings": { ... updated settings ... }
}
| Setting |
Range |
Default |
Description |
commit_merge_ms |
0 → 10000 |
2500 |
Milliseconds to buffer voice-activity-detection (VAD) commits before translating; merges short fragments into complete phrases. |
stability_timeout_ms |
0 → 5000 |
2000 |
Milliseconds to wait for stable partial text (unchanged for this duration) before dispatching for translation; fallback when VAD commits don’t fire. |
tts_segment_pause_ms |
0 → 1000 |
0 |
Pause inserted between consecutive TTS audio segments on the frontend; allows speaker to breathe between sentences. |
max_accumulation_ms |
0 → 30000 |
8000 |
Maximum time to accumulate new words during continuous speech before forcing translation; prevents long speech from batching indefinitely. |
vad_threshold |
0.0 → 1.0 |
0.5 |
Voice-activity-detection sensitivity; higher = stricter noise filtering, lower = catches more audio but accepts more false positives. |
vad_silence_threshold_secs |
0.5 → 3.0 |
1.5 |
Seconds of silence before VAD automatically commits the current phrase and triggers translation. |
min_speech_duration_ms |
50 → 500 |
100 |
Ignore audio shorter than this duration (suppresses clicks, pops, very brief utterances). |
min_silence_duration_ms |
50 → 500 |
100 |
Minimum gap (silence) required between speech chunks; prevents word fragments separated by brief pauses from being rejoined. |
flush_on_sentence_boundary |
true | false |
true |
When true, dispatch complete sentences at (.?!;) boundaries instead of waiting for longer commits; provides snappier output at risk of incomplete sentences. |
min_chars_before_dispatch |
10 → 500 |
40 |
Minimum characters in a transcript chunk before it is sent for translation; prevents tiny fragments from consuming translation quota. |
Video Call Settings
Separate STT/TTS configuration for low-latency video translation calls. Uses faster TTS model (eleven_flash_v2_5) and tighter stability windows.
GET /admin/video-settings
Returns current video call TTS/STT settings.
Response:
{
"stability_ms": 500,
"commit_merge_ms": 50,
"translation_provider": "claude"
}
POST /admin/video-settings
Update video call settings.
Request Body:
{
"stability_ms": 500,
"commit_merge_ms": 50,
"translation_provider": "claude"
}
Response:
{
"stability_ms": 500,
"commit_merge_ms": 50,
"translation_provider": "claude"
}
| Setting |
Range |
Default |
Description |
stability_ms |
100 → 2000 |
500 |
Milliseconds to wait for stable partial before dispatching; shorter window for low-latency calls. |
commit_merge_ms |
0 → 500 |
50 |
VAD commit merge delay; very short (50ms) for snappy real-time response in video. |
translation_provider |
libretranslate | claude | deepl | google |
claude |
Translation provider for video calls; can differ from global broadcast setting. |
TTS Pipeline Configuration (application.yaml)
Buffering and pacing parameters for the TTS worker pool. Configured in application.yaml under tts_pipeline.
| Setting |
Range |
Default |
Description |
initial_buffer_segments |
1 → 10 |
1 |
Number of translated segments to buffer before TTS playback starts; higher = longer delay before audio begins but fewer playback stalls. |
low_water_hold_ms |
0 → 5000 |
1500 |
Hold audio emit until the next segment is buffered (for up to this many ms); prevents frontend from running out of queued audio. Set to 0 to disable. |
audio_lag_segments |
0 → 10 |
2 |
How many segments TTS audio should trail behind live translation text; keeps audio playback behind the scrolling transcript. |
audio_lag_timeout_ms |
1000 → 15000 |
8000 |
Maximum time to wait for audio lag buffer to fill before emitting anyway; prevents audio gaps during speaker pauses. |
ElevenLabs Configuration (application.yaml)
Core ElevenLabs service settings. Configure in application.yaml under elevenlabs.
| Setting |
Range |
Default |
Description |
api_key |
string |
${ELEVENLABS_API_KEY} |
ElevenLabs API key (required); set via environment variable in production. |
default_voice_id |
string |
kxj9qk6u5PfI0ITgJwO0 |
Default voice ID for TTS when none is explicitly selected; must be a valid ElevenLabs voice ID. |
tts_model |
string |
eleven_multilingual_v2 |
TTS model identifier; options: eleven_multilingual_v2, eleven_turbo_v2, eleven_flash_v2_5 (video calls). |
stt_model |
string |
scribe_v2_realtime |
STT (speech-to-text) model; scribe_v2_realtime is the low-latency WebSocket API. |
Notes
- Runtime Updates: Changes made via
POST /admin/tts-settings are persisted to Redis and apply immediately to all new broadcasts. In-flight broadcasts complete with their original settings.
- Voice Stability vs. Similarity: Higher stability produces more consistent pronunciation but may sound less natural. Higher similarity boost makes output match the voice sample more closely but may introduce artifacts.
- STT Timing Trade-offs: Short commit merge and stability windows produce snappier output but risk splitting sentences. Long windows batch more text together but introduce latency before translation starts.
- Sentence Boundary Flushing: When enabled, transcripts dispatch at (.?!;) punctuation, making output more granular but potentially incomplete (e.g., a sentence without a period). Disable to wait for full VAD commits instead.
- Video vs. Broadcast: Video calls use
video-settings (100–500ms windows, flash TTS model) for low-latency interaction. Broadcasts use global tts-settings with longer buffers for quality.
STT Timing Configuration
These settings control how long the speech-to-text engine waits before dispatching audio for translation. Adjust to balance responsiveness (shorter waits = snappier feedback) against reliability (longer waits = fewer false fragments).
Settings Table
| Setting |
Default |
Description |
commit_merge_ms |
2500 |
Buffer VAD commits (ms) before translating — merges speech fragments into larger chunks. |
stability_timeout_ms |
2000 |
Wait for stable partial text (ms) before translating when no VAD commit fires. |
tts_segment_pause_ms |
0 |
Pause between TTS audio segments (ms) — sent to frontend for playback pacing. |
max_accumulation_ms |
8000 |
Max time to accumulate words during continuous speech before force-dispatching (ms). |
vad_threshold |
0.5 |
Voice Activity Detection sensitivity (0–1) — higher = stricter noise filtering. |
vad_silence_threshold_secs |
1.5 |
Silence duration (seconds) before VAD triggers a commit. |
min_speech_duration_ms |
100 |
Ignore speech shorter than this (ms) — filters brief noise spikes. |
min_silence_duration_ms |
100 |
Minimum silence gap (ms) to distinguish separate speech segments. |
flush_on_sentence_boundary |
true |
When true, dispatch at sentence boundaries (.?!) instead of waiting for longer pauses. |
min_chars_before_dispatch |
40 |
Minimum characters in a chunk before translation — prevents tiny fragments. |
Tuning Guide
- Snappier response: Lower
commit_merge_ms (1000–1500), stability_timeout_ms (800–1200), and max_accumulation_ms (4000–5000). Trade-off: more fragments → TTS queue churn & potential audio gaps.
- Smoother audio playback: Raise
max_accumulation_ms (10000–12000) and min_chars_before_dispatch (50–80) to batch larger chunks. TTS has more time between segments — audio queue stays deeper.
- Reduce false positives: Increase
vad_threshold (0.6–0.8) to reject background noise; raise min_speech_duration_ms (150–200) to ignore brief clicks.
- Faster VAD commits: Lower
vad_silence_threshold_secs (0.8–1.0) so pauses trigger sooner. Caution: may split sentences mid-clause if speaker hesitates.
- Disable sentence boundaries: Set
flush_on_sentence_boundary to false to dispatch based on character/word thresholds alone (useful for languages without clear punctuation).
API
Retrieve and update STT timing settings via REST:
GET /admin/stt-timing
Response: { "settings": { "commit_merge_ms": 2500, "stability_timeout_ms": 2000, … } }
POST /admin/stt-timing
Request body: { "max_accumulation_ms": 6000, "tts_segment_pause_ms": 200 }
Response: { "settings": { …updated fields… } }
All settings are optional in the POST request — only provided fields are updated.
Notes
- Startup default: Settings load from
application.yaml, then from Redis if persisted.
- Runtime changes: Updates via the admin UI take effect immediately on all connected clients.
- TTS pacing:
tts_segment_pause_ms is a UI hint; the frontend's audio player waits this long between segment playbacks.
- Accumulation timer: Runs continuously during speech to ensure translation happens every N ms, even when VAD & stability timers don't fire (e.g. long sermons with minimal pauses).
Authentication: All endpoints require JWT cookie authentication (adminAuth middleware). Users must have is_admin=true OR hold a role with relevant permissions. Permissions are checked per-endpoint via requirePermission(...).
API Keys
Retrieve all configured API keys and their status (configured/missing).
Update one or more API keys. Valid names: elevenlabs, anthropic, deepl, libretranslate, google, youtube.
Body: { elevenlabs?: string; anthropic?: string; deepl?: string; libretranslate?: string; google?: string; youtube?: string }
Voice Management
Scan and list all available ElevenLabs voices with their categories and preview URLs.
Get the list of voice IDs allowed for viewers to select (null = all voices allowed).
Restrict viewer access to a specific set of voice IDs. Broadcasts to all connected clients.
Body: { voiceIds: string[] }
Feature Flags
Get all feature flags merged from YAML config + Redis overrides.
Get the current value of a single feature flag.
Set a feature flag value. When auto_pause_songs is disabled, releases any held auto-pause. Broadcasts to all clients via Socket.IO.
TTS & STT Settings
Retrieve current TTS voice settings (stability, similarity_boost, style, speed, speaker_boost).
Update TTS settings. Only provided fields are changed; others retain current values.
Body: { stability?: number; similarity_boost?: number; style?: number; speed?: number; use_speaker_boost?: boolean }
Get STT timing settings (commit merge delay, stability timeout, VAD parameters, min_chars_before_dispatch).
Update STT timing settings. Controls how long to buffer before translation and sentence-boundary flushing.
Body: { commit_merge_ms?: number; stability_timeout_ms?: number; max_accumulation_ms?: number; vad_threshold?: number; vad_silence_threshold_secs?: number; min_speech_duration_ms?: number; min_silence_duration_ms?: number; flush_on_sentence_boundary?: boolean; min_chars_before_dispatch?: number }
Generate a TTS audio preview for a text sample. Returns audio/mpeg or audio/L16 (PCM) based on format parameter.
Body: { text: string; voiceId?: string; format?: 'mp3' | 'pcm' }
Languages
Get the currently active language pair (source → target, e.g. ['en', 'ru']).
Set the active language pair. Must be exactly 2 language codes. Broadcasts update to all viewers in real-time.
Body: { languages: [string, string] }
Get the pool of language codes viewers can select from (admin-curated list).
Set the available language pool. Broadcasts both the pool and current active pair to all clients.
Body: { languages: string[] }
Translation Provider
Get the currently active translation provider and list of available options (google, deepl, claude, libretranslate).
Switch the active translation provider at runtime. Must be one of: google, deepl, claude, libretranslate.
Body: { provider: 'google' | 'deepl' | 'claude' | 'libretranslate' }
Get the currently selected Claude translation model and list of available models.
Select a specific Claude model for translation (e.g. claude-opus-4-1, claude-sonnet-4).
Audio Device
Get the admin-selected audio input device (overrides viewer's local choice). Returns deviceId and label.
Force all viewers to use a specific audio input device. Broadcasts the selection to all connected clients via Socket.IO.
Body: { deviceId?: string; label?: string }
Audio Stream URL
Get the saved HTTP audio stream URL (e.g., Icecast endpoint for church broadcasts).
Save an HTTP(S) audio stream URL. Used by the 'stream' broadcast source. Empty string clears the saved URL.
Video Settings
Get video call translation settings (stability_ms, commit_merge_ms, translation_provider).
Update video call STT/TTS settings (separate from main live-translation settings).
Body: { stability_ms?: number; commit_merge_ms?: number; translation_provider?: 'libretranslate' | 'claude' | 'deepl' | 'google' }
Sermon Generation
Generate a sermon snippet via Gemini Flash. Returns poetic biblical text (Russian, Ukrainian, or English).
Body: { apiKey?: string; language?: 'ru' | 'uk' | 'en'; sentences?: number }
Voice Training & Cloning
Create an instant voice clone from browser mic recordings. Accepts base64-encoded audio blobs (webm, mp3, etc.).
Body: { name: string; clips: string[]; mimeType?: string }
Create an instant voice clone from a YouTube video. Downloads and extracts N×30s clips via yt-dlp + ffmpeg.
Body: { name: string; youtubeUrl: string; clipCount?: number; startOffset?: number }
Broadcast Schedule
Get the list of scheduled broadcast events with optional auto-start configuration.
Update the broadcast schedule. Each event includes id, title, datetime (ISO 8601), and optional source config for Phase 2 auto-start.
Body: { events: Array<{ id: string; title: string; datetime: string; description?: string; source?: 'mic'|'youtube'|'biblical'|'remote'|'stream'; voiceId?: string; youtubeUrl?: string; biblicalPrompt?: string; agentId?: string; allowedLanguages?: [string, string] }> }
Hallucination Monitoring
Get hallucination detection statistics (count, recent examples, filter reasons).
Clear the hallucination detection log.
Custom Fillers
Get the list of custom filler words (speaker-specific utterances like "uh", "um", "угу") to strip before translation.
Set custom filler words. These are stripped from transcripts on-the-fly to avoid translation API calls and garbled output.
Body: { words: string[] }
Translation Log
Get the translation history log (original text, translations, timings, detected language, provider used).
Clear the translation log.
Video Room Moderation
List all active video call rooms with participant details (language, display name, connection status). Participant tokens are stripped.
Force-close a video call room by code. Broadcasts disconnect notification to all participants.
Queue Monitoring
Get real-time snapshot of translation pipeline queue depths and Redis Stream stats.
Session History
Retrieve broadcast session history from PostgreSQL (list of completed broadcasts with timestamps and sources).
Get detailed transcript log for a single broadcast session (original text, translations, detected language, timings).
Export a session transcript in CSV, TXT, or JSON format (default: JSON). Query param: ?format=csv|txt|json.
User Management
List all users. Requires user_management permission. Password hashes and avatar data are stripped from response.
Update a user's role assignment. Requires user_management permission. Supports both legacy single-role and new multi-role formats.
Body: { isAdmin?: boolean; roleId?: string | null; roleIds?: string[] }
Reset a user's password. Requires user_management permission. Password must be at least 6 characters.
Body: { password: string }
Delete a user account. Requires user_management permission. Cannot delete your own account.
Permissions & Roles
List all available permissions. Requires user_management permission.
List all roles with their permission assignments. Requires user_management permission.
Create a new role with a name and list of permissions. Requires user_management permission. Role names must be unique.
Body: { name: string; permissions: Permission[] }
Update a role's name and/or permissions. Requires user_management permission.
Body: { name: string; permissions: Permission[] }
Delete a role. Requires user_management permission.
YouTube
Get the currently configured YouTube channel ID and whether it was loaded from environment or persisted in Redis.
Set the YouTube channel ID for live stream lookups. Persisted to Redis; overrides environment variable.
Body: { channelId: string }
Find live streams on a YouTube channel. Uses YouTube Data API if key is configured, falls back to yt-dlp. Query param: ?channelId=UCxxxxxx (optional; uses saved channel ID if not provided).
Public Endpoints (No Auth Required)
Note: This endpoint appears to lack adminAuth middleware in the source code. Returns the configured Anthropic API key (potential security issue — should require authentication).
SDK
Uses the official @elevenlabs/elevenlabs-js SDK (v2). The client is lazy-loaded on first use.
Speech-to-Text (Scribe v2 Realtime)
Connects via native WebSocket to wss://api.elevenlabs.io/v1/speech-to-text/realtime. Handles:
- VAD-based commit buffering with configurable merge window
- Stability timeout fallback for stalled VAD
- Text validation (EN/RU/UK character regex filtering)
- Partial and final transcript emission
Text-to-Speech
Uses client.textToSpeech.stream() with the eleven_multilingual_v2 model. Audio is collected into a Buffer and emitted as base64 MP3.
Voice Management
client.voices.getAll() — fetches all voices from account
- Admin can filter which voices are available to viewers
- Voice cloning via IVC API (from recordings or YouTube)
Key File
backend/src/services/elevenlabs.service.ts
Provider Details
Google Translate
Google Cloud Translation API v2. Fast (~200ms), deterministic, and reliable. Requires GOOGLE_TRANSLATE_API_KEY with the Cloud Translation API enabled in Google Cloud Console. Ensure the API key has no HTTP referrer restrictions (server-side requests have no referrer).
File: backend/src/services/google-translate.service.ts
LibreTranslate
Self-hosted in Docker. No API key required by default. Provides language detection and translation via REST API.
File: backend/src/services/libretranslate.service.ts
DeepL
Premium translation API. Auto-detects free vs. paid endpoint based on the API key format.
File: backend/src/services/deepl.service.ts
Claude (Anthropic)
AI-powered translation using claude-haiku-4-5 for speed. Includes language detection and auto-flip logic.
File: backend/src/services/claude-translate.service.ts
Routing
Provider routing is handled by backend/src/services/translation.provider.ts:
- Try admin-selected primary provider
- On failure, try configured fallback provider
- LibreTranslate is always the last-resort fallback
Connection
Uses ioredis with automatic retry strategy. Falls back to in-memory/YAML defaults if Redis is unavailable.
Key Patterns
| Pattern | Example | Purpose |
flag:<name> |
flag:youtube_input |
Feature flag boolean values |
setting:<name> |
setting:tts_settings |
JSON settings objects |
Key File
backend/src/services/redis.service.ts
Local Development
Use docker-compose.local.yml for Redis and LibreTranslate only (backend/frontend run natively):
docker compose -f docker-compose.local.yml up -d
Production
Use docker-compose.yml for all services:
docker compose up -d --build
Services
| Service | Image | Port | Notes |
| frontend |
node:24-alpine + Nginx |
80 (exposed) |
Serves React build, proxies API/WS to backend |
| backend |
node:24-alpine |
3001 (internal) |
Express + Socket.io server |
| redis |
redis:7-alpine |
6379 (internal) |
Feature flags and settings store |
| libretranslate |
libretranslate/libretranslate |
5000 (internal) |
Self-hosted translation engine |
Configuration
ELEVENLABS_API_KEY=sk-your-production-key
ADMIN_PASSWORD=strong-secure-password
FRONTEND_URL=https://translate.example.com
APP_ENV=prod
REDIS_PASSWORD=redis-secret
Deploy
docker compose up -d --build
Reverse Proxy
When running behind Nginx or another reverse proxy:
- Set
LISTEN_PORT in .env (e.g., 8080)
- Proxy pass to
localhost:8080
- Important: Ensure WebSocket upgrades are forwarded for the
/socket.io/ path
server {
listen 443 ssl;
server_name translate.example.com;
location / {
proxy_pass http://localhost:8080;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
}
}
Monitoring
# Check all services
docker compose ps
# View backend logs
docker compose logs -f backend
# Health check
curl http://localhost:3001/api/health