Skip to content

About

Jev VAD: client-side turn detection for Agora Conversational AI — manual SoS/EoS and semantic barge-in judged by TypeSafe Jev

Resources

Stars

20 stars

Watchers

0 watching

Forks

Repository files navigation

Jev VAD — client‑side turn detection for Agora ConvoAI

Agora Conversational AI agent where Agora's turn detection is fully switched off and every turn decision — start of speech, end of speech, barge‑in — is made client‑side. The design is open fast, decide late: a dumb acoustic gate opens the turn in ~110 ms, and TypeSafe Jev decides everything that matters from the text the STT streams back — is the user done, are they talking to the agent at all, are they really interrupting or just saying "uh‑huh".

Agora is reduced to a pipe (RTC + STT + LLM + TTS). The app tells it exactly when a user turn starts (manualSOS), when it ends (manualEOS), and when to stop talking (interrupt).

 mic ──▶ SNR gate (~110 ms, drops single impulsive bursts) ──▶ manualSOS
                                                                  │
                                                     Agora STT streams text
                                                                  ▼
 mic quiet clock ─┐          ┌──────────────── one Jev call ────────────────┐
 role‑tagged hist ┼──▶       │ turn_complete · holding_floor ·             │
 text trajectory ─┤          │ addressed_to_agent · taking_floor           │
 asr final/stable ┘          └──────────────────────┬────────────────────────┘
                                                    ├─ complete & addressed & quiet ──▶ manualEOS
                                                    ├─ not addressed / backchannel ───▶ hold (no EoS)
                                                    └─ agent talking & taking_floor ─▶ interrupt()

Why

Server VADs decide on silence. Humans don't. "Yeah… mmm, I'm thinking about it" has a pause after yeah that any silence VAD fires on. Jev reads the transcript with the assistant's last question in front of it and answers a different question: has this person handed me the floor? Silence becomes one input among several instead of the trigger.

The same trick covers two things a silence VAD can't do at all: addressee detection ("hey Mark, dinner's ready" opens a turn but never reaches the LLM) and backchannel rejection ("uh‑huh", "right" while the agent talks don't cut it off; "wait, stop" does).

Why not judge the audio?

v1 asked Jev to classify a 220 ms window of 16 acoustic features before opening the turn. Two problems:

  • manualSOS gates what the STT hears. Audio before it is discarded, so 220 ms hold + ~300 ms Jev + RTM clipped the first word of every utterance.
  • The STT is already an excellent speech/non‑speech classifier — a cough, a chair, a keyboard produce no text. Opening a turn on them costs nothing as long as you never send EoS for an empty turn (see spike findings below).

So the gate opens immediately and Jev works where it is strong: text in conversational context.

Architecture

Layer Where Role
Acoustic gate web/src/lib/mic-energy.ts Web Audio analyser. Tracks a noise floor, gates on SNR, nominates after 110 ms of sustained energy. Drops obvious single impulsive bursts locally. Stricter gate + 350 ms refractory while the agent talks (TTS echo). Exposes quietForMs() — a mic‑based silence clock that runs even while a turn is open.
Turn state machine web/src/lib/turn-machine.ts Pure, framework‑free. dispatch(ctx, event, config) → {ctx, effects}. Phases: armed → opening → open_empty → open_text → judging → waiting / confirming / hold → ending. Unit‑testable without a browser.
Controller web/src/hooks/useTurnController.ts Hosts the machine: timers, the mic gate, /api/judgeTurn, and manualSOS / manualEOS / interrupt via agora-agent-client-toolkit. Exposes window.__jevvad for diagnostics.
Live config web/src/lib/vad-config.ts · VadConfigPanel.tsx Every knob is runtime‑tunable from the Tune drawer during a call; persisted in localStorage. Jev thresholds ride along with each judgeTurn request, so no server restart.
Jev judgment server/src/jev_eos.py One TypeSafe system_one call with four nouls (see below). Returns nouls, thresholds and token usage.
Agent config server/src/agent.py start_of_speech.mode: manual, end_of_speech.mode: manual, interruption.enable: false, silence_config.timeout_ms: 0, audio_scenario: chorus.
UI web/src/components/JevVadPanel.tsx Phase, four live score bars with threshold markers, hold reason, interrupted chip, per‑session token cost.

Agora configuration

turn_detection={
    "mode": "default",
    "config": {
        "start_of_speech": {"mode": "manual"},
        "end_of_speech":   {"mode": "manual"},
    },
},
interruption={"enable": False, "disabled_config": {"strategy": "append"}},
parameters={"audio_scenario": "chorus", "data_channel": "rtm", "silence_config": {"timeout_ms": 0}, ...},
  • manual SoS: the server only counts audio toward a turn after manualSOS.
  • manual EoS: the turn is submitted to the LLM only on manualEOS.
  • interruption disabled: voice never stops the agent on its own; only our explicit interrupt() does.
  • silence timeout 0: no "are you still there?" prompts — the server never owns idle time.

1. Start of speech — the gate

watchMicEnergy runs for the whole session. Noise floor: fast attack downward, very slow upward, frozen while a turn is open. Gate = max(0.006, floor × 3) (≈ 9.5 dB), or floor × 5 while the agent is speaking. 60 ms dip grace. After 110 ms above the gate the window is summarised; if it is an obvious single burst (peak_position < 0.35 && decay_ratio < 0.3 && continuity < 0.6) it is dropped, otherwise manualSOS fires immediately.

The transcript present at that moment is recorded as a baseline so the previous turn's text can't be judged as this turn's. If no text arrives within 4 s the open is marked a false open: the turn stays open, nothing is sent, and the user's next words land in it un‑clipped.

2. The judgment — one call, four nouls

Every meaningful transcript change (debounced 150 ms) and every silence re‑check makes one POST /api/judgeTurn. Jev sees only text plus client timing:

{
  "setting": "Voice call between a person and an AI assistant named Ada… may be in a room with other people…",
  "transcript_notes": "Live partial from Deepgram. Auto‑punctuation is not evidence of completion…",
  "timing_notes": "silence_since_last_word_ms comes from the mic; asr_final/stable_word_ratio from the recognizer…",
  "conversation_history": [{"role": "assistant", "text": "…"}, {"role": "user", "text": "…"}],
  "assistant_last_message": "Which platform are you building on?",
  "assistant_state": "listening",
  "current_user_transcript": "Web and mobile.",
  "transcript_trajectory": [{"text": "Web", "ms_ago": 900}, {"text": "Web and", "ms_ago": 500}],
  "turn_elapsed_ms": 1400,
  "silence_since_last_word_ms": 800,
  "asr_final": true,
  "stable_word_ratio": 1.0
}
noul question threshold
turn_complete Has the user handed over the floor? eosThreshold 0.72 at 0 ms quiet, relaxing linearly to eosThresholdFloor 0.55 over eosDecayMs 3 s of mic silence
holding_floor Are the latest words asking for time ("let me think", "hmm")? A later "please continue" releases it. holdingFloorMin 0.5 → patient mode, no bar relaxation
addressed_to_agent Is this for the assistant, or for someone else in the room / a phone / a TV? Judged by the most recent part if mixed. addresseeThreshold 0.5
taking_floor (agent speaking only) Deliberate interruption, or a backchannel / echo fragment? Gets assistant_current_speech. bargeInThreshold 0.5

The prompts are explicit that Jev only sees text: prosody, hums, breaths are invisible unless Deepgram rendered them as words. Silence therefore comes from the mic, not from a static transcript. Worked examples anchor short answers ("I don't know." after a question → complete), dangling connectors ("…but" → not), lead‑ins ("I have a question." → not), and ASR periods mid‑thought ("So the thing is." → not).

3. End of speech — the decision loop

judgeResult
├─ agent speaking & taking_floor ≥ thr → interrupt()          (then continue below)
├─ agent speaking & taking_floor < thr → HOLD (backchannel)   agent keeps talking
├─ addressed_to_agent < thr          → HOLD (side‑talk)      no EoS, re‑judge on new text
├─ turn_complete ≥ bar(quiet)       bar = 0.72 → 0.55 over 3 s of mic silence
│    │                               (bar stays 0.72 while holding_floor ≥ 0.5)
│    └─ wait until mic quiet ≥ confirm(turn_complete)   350–1000 ms, shorter when confident
│         └─ manualEOS
└─ not yet
     ├─ mic quiet ≥ cap → manualEOS (timeout)   cap = 5 s, or 12 s if holding_floor ≥ 0.5
     └─ else wait(turn_complete)                700–2500 ms, longer the further below the bar
          └─ when the mic has truly been quiet that long → judge(text, silence_ms)

Jev's score is the decision; the client only moves the bar with silence. Measured on real sessions, complete short questions land at 0.67–0.70 and barely drift with silence, so a fixed 0.72 stalled them until the 5 s cap. With the relaxing bar they close after ~1.5–2 s of quiet, while an "I have a few questions… pricing." at 0.50–0.54 still waits (floor 0.55).

New transcript text at any point supersedes the loop (sequence number + AbortController) — except during barge‑in, where in‑flight judgments are kept and a stale "interrupt" verdict on a prefix of the current text is acted on immediately. Empty turns never fire EoS.

4. Barge‑in (two stage)

  • Stage 1 — acoustic. Onset while the agent is speaking/thinking → manualSOS only, with the stricter echo gate. The agent keeps talking; STT starts listening.
  • Stage 2 — semantic. First words arrive → Jev taking_floor. "Uh‑huh." → 0.03, hold. "Uh‑huh. Wait." → 0.86, interrupt(). The turn then continues to a normal EoS. While the agent talks there is no debounce (bargeInDebounceMs 0) and every partial gets its own judgment; earlier ones are not cancelled, so the first partial that reads as a floor grab interrupts as soon as its verdict lands (~300 ms after the words appear).

A cough that passes stage 1 never reaches stage 2 because the STT emits nothing.

Live tuning

The Tune button in the call header opens a drawer with every knob, grouped:

  • Gate — hold (ms), gate ×floor idle / agent‑talking, agent refractory, drop impulsive bursts.
  • Timing — empty turn, judge debounce, barge‑in debounce, confirm quiet min/max, re‑check min/max, silence cap normal / thinking, holding floor ≥.
  • Judgment — complete ≥, complete ≥ (after silence), relax over, to agent ≥, interrupting ≥, side‑talk hold on/off, semantic barge‑in on/off (off = any transcribed words interrupt).

Changes apply on the next event (gate values on the next onset, timing on the next judgment, thresholds on the next verdict — the client sends them with each request), persist in localStorage under jevvad.config.v1, and Reset restores defaults. Env JEV_*_THRESHOLD values remain the server fallback when a request carries no overrides.

Spike findings (server behaviour under manual mode)

Measured against the live agent with a headless Chromium and a WAV file as the microphone (--use-file-for-fake-audio-capture), driving window.__jevvad.sos()/eos():

scenario result
manualSOS → 3 s silence → manualEOS with empty ASR LLM runs. Agent replies "It looks like your message didn't come through." → never EoS an empty turn.
manualSOS → silence, wait for server AGENT_MANUAL_EOS No auto‑close within 120 s, no reply, agent stays connected. An open empty turn is harmless.
manualSOS via the gate on "Hey Mark, dinner is ready in five minutes…" Gate opened at 114 ms, STT transcribed, addressed_to_agent = 0.03–0.09 → hold. No reply within 110 s.
"uh‑huh" then "wait, wait, stop, what is Agora?" during an agent answer Backchannel → hold (taking_floor 0.03). "Wait." → interrupt() (0.86). Turn completed → answered "Agora is…".
"Can you explain in detail, step by step, how…" through the 110 ms gate First transcribed partial was "Can" — no clipped first word.

Consequence for side‑talk: the words stay in the open turn, so when the user later addresses the agent the LLM receives both. ADA_PROMPT tells the agent to ignore parts clearly meant for someone else, and the addressee prompt judges mixed turns by their most recent part.

Cost

Measured on this prompt set (usage is returned by the SDK and surfaced in the panel and the backend log):

call input tok output tok latency (p50 / max)
judgeTurn, agent idle (3 nouls) ~2,100 ~57 ~275 ms / ~500 ms
judgeTurn, agent speaking (4 nouls) ~2,450 ~74 ~300 ms / ~325 ms

Per user turn in practice: 3–6 judgments (transcript partials plus silence re‑checks) ≈ 7–13k tokens. No Jev call is made at SoS any more, and none for false opens.

Levers if you need to cut cost: raise judge debounce (fewer partial judgments), cap conversation_history at 4 turns, raise re‑check min so fewer silence re‑checks fire.

Running it

bun run setup        # venv + deps; writes server/.env via the Agora CLI
bun run dev          # backend :8000 + web :3000

server/.env:

AGORA_APP_ID=…
AGORA_APP_CERTIFICATE=…
TYPESAFE_API_KEY=…
JEV_EOS_THRESHOLD=0.72
JEV_ADDRESSEE_THRESHOLD=0.5
JEV_BARGE_IN_THRESHOLD=0.5

Expose publicly (mic needs HTTPS off‑localhost): ngrok http 3000. The Next dev server already allow‑lists *.ngrok* origins.

Backend endpoint: POST /judgeTurn {transcript, recent_turns, silence_ms, trajectory, turn_elapsed_ms, asr_final, stable_word_ratio, agent_state, assistant_current_speech, thresholds?} → {turn_complete, holding_floor, addressed_to_agent, taking_floor, should_end, should_interrupt, thresholds, usage}. The Next app proxies it under /api/judgeTurn.

Diagnostics in DevTools: __jevvad.state() (machine context), __jevvad.log (onsets, verdicts, phase changes, server acks), __jevvad.sos() / eos() / interrupt().

Known limitations

  • No way to discard an open turn. Side‑talk and backchannels stay in the turn until the user addresses the agent; mitigated by the system prompt, not eliminated.
  • Echo. Browser AEC handles most TTS bleed; the stricter agent‑speaking gate and the semantic stage‑2 check handle the rest. Loud speakers can still leak transcribable words.
  • Single speaker. No diarisation; anyone at the mic is "the user". Addressee detection is semantic, not acoustic.
  • Jev is text‑blind to prosody. A rising‑tone "yeah?" and a flat "yeah." look identical. The mic‑quiet confirmation window is the mitigation.

Stack

Next.js 16 · React · agora-agent-client-toolkit · agora-agent-uikit · FastAPI · agora-agents Python SDK · typesafe-sdk (Jev) · Deepgram STT · OpenAI · MiniMax TTS.

About

Jev VAD: client-side turn detection for Agora Conversational AI — manual SoS/EoS and semantic barge-in judged by TypeSafe Jev

Resources

Stars

20 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages