How to build semantic turn detection with Jev on top of VAD
Use VAD to detect silence, then ask Jev, TypeSafe's decision model, whether the user has finished. VAD alone ends the turn after a fixed silence, so a user who pauses mid-thought gets cut into two turns, and raising the threshold for everyone makes short answers feel slow. At taritas we combine three signals: VAD with an adaptive noise floor, trailing words such as 'and' or 'but', and one Jev yes/no question after 350 ms of silence: has the speaker finished their thought? We send the agent's last message with the user's words, because 'Yes.' is only complete after a question. Jev returns a probability in about 0.4 s, inside the pause, and that probability sets the wait: 1.7 s for a finished sentence, 2.8 s when unsure, and 4.5 s for one that trails off. Users can also end the turn at once.
Published · Updated · Supreet Tare
All names, numbers, and identifiers in this post are anonymized. The patterns are real.
People pause when they think. “I tried calling him yesterday, and…” is often followed by a second of silence before the sentence continues. A voice agent that feels natural waits for the rest.
This post is the turn detection setup we use in production on a bilingual voice coaching agent (English and a second language). It covers the three signals, the Jev check that handles the semantic part, the wait times we use, the code, and a checklist you can reuse. For why a fixed silence timeout fails in the first place, and what it broke on a production phone line, see our earlier post on semantic turn detection.
VAD and turn detection: four terms in one table
| Term | What it answers | Typical output |
|---|---|---|
| VAD (voice activity detection) | Is someone speaking right now? | Speech or silence, frame by frame |
| Endpointing | When should the transcriber finalize the text? | A “commit” after a fixed silence, often 0.5 to 1 s |
| Turn detection | Has the user finished their turn? | End of turn, yes or no |
| Semantic turn detection | Is the thought complete, given the words and the context? | A probability, 0 to 1 (from Jev, in our build) |
Most realtime stacks ship with VAD and endpointing. Turn detection is the layer you add on top.
Why VAD alone is not enough
With VAD only, the turn ends when the silence reaches a fixed threshold. Ours started at 0.6 s, the transcriber’s endpointing default.
That works well for short commands. In open conversation, people pause mid-thought, and a fixed threshold splits one thought into two turns. Raising the threshold for everyone makes short answers like “Yes.” feel slow.
The fix is to keep VAD as the trigger and let the words decide how long to wait.
The three-signal design
1. Acoustic (VAD). We measure the microphone level against a noise floor that adapts to the room. A noise floor is the background level of the room, so “silence” means quiet for this room, not a fixed decibel value. Any new sound resets the clock.
2. Lexical. Some endings mean “more is coming”: a comma, ”…”, or a trailing word like “and”, “but”, “because”. These are cheap string rules with one word list per language.
3. Semantic (Jev). After 350 ms of silence, we ask Jev one question: has the speaker finished what they wanted to say? We send the words so far and the agent’s last message, because context changes the answer. “Yes.” is a complete reply to a yes/no question. “I think the thing is” is incomplete after anything.
Why Jev fits turn detection
Jev is TypeSafe’s typed decision model, released in early access in September 2026. It answers typed questions with probabilities instead of generating text. That shape suits turn detection well:
- One question, one number. “Has the speaker finished?” returns a probability that goes straight into a threshold, with no text to parse.
- Fast enough to run inside a pause. A Jev check took about 0.4 s in our build.
- Context-aware. Jev reads the agent’s last message alongside the user’s words, which VAD and word rules cannot do.
- Works across languages. The same Jev question worked in both languages we support.
Recommended starting values
These are our production defaults. They are a good place to start, and worth tuning with your own users.
| Setting | Value |
|---|---|
| Jev check starts after | 350 ms of silence |
| Jev check latency | about 0.4 s |
| Jev “finished” threshold | probability 0.75 or more |
| Wait if finished | 1.7 s |
| Wait if unsure | 2.8 s |
| Wait if trailing off (“and…”, a comma) | 4.5 s |
| Transcriber flush | 0.4 s before the decision |
| Re-check | only if the words change |
| User control | adjustable wait, plus Send button or tap on mic to end the turn at once |
Implementation with Jev
The lexical rule checks whether the text ends in a way that promises more:
TRAILING_WORDS = {
"en": {"and", "but", "because"},
"lang2": {...}, # the same connectors in your second language
}
def ends_mid_thought(text: str, lang: str) -> bool:
t = text.rstrip()
if t.endswith((",", "...")):
return True
last = t.split()[-1].lower().strip(".,!?") if t else ""
return last in TRAILING_WORDS[lang]
The Jev turn detection check is one typed question. The request, simplified, looks like this:
state = {
"agent_last_message": agent_last_message,
"user_words_so_far": joined_transcript,
}
questions = {
"finished": {
"type": "yes_no",
"question": "Has the speaker finished what they wanted to say?",
},
}
p_finished = jev_decide(state, questions)["finished"] # probability, 0 to 1
The wait policy maps the lexical rule and the Jev probability to a wait time:
def wait_seconds(p_finished: float, text: str, lang: str) -> float:
if ends_mid_thought(text, lang):
return 4.5
if p_finished >= 0.75:
return 1.7
return 2.8
The transcript arrives in segments, and a segment boundary can look like a full stop. Join segments across pause marks before you ask Jev:
segment 1: "I was going to tell her and..."
segment 2: "When I got there"
naive join: "I was going to tell her and... When I got there"
clean join: "I was going to tell her and when I got there"
With the naive join, Jev scored a clearly unfinished thought at 0.36. With the clean join, it read the sentence as one unfinished thought.
Results
- Accuracy: Jev classified 16 of 16 test phrases correctly across both languages. Finished thoughts scored 0.82 or higher, and mid-thought phrases 0.06 or lower. The set is small, but the wide gap between the two groups is what makes a single threshold dependable.
- Latency: about 0.4 s per Jev check, inside the silence the agent is already waiting through.
- End to end: in browser tests with synthetic speech, a sentence with a 2.5 s pause in the middle was held as one turn, in both languages.
Turn detection checklist
- VAD with an adaptive noise floor, and any new sound resets the clock.
- One trailing-word list per language, plus comma and ”…” rules.
- Jev check starts after about 350 ms of silence.
- The Jev check gets the agent’s last message, not only the user’s words.
- Ask once per pause, and again only if the words change.
- Flush held-back transcriber words before the decision.
- Join transcript segments across pause marks.
- Map the Jev probability to a wait time, not straight to “reply now”.
- Give users an adjustable wait and a way to end the turn at once.
- Test with sentences that contain a long mid-sentence pause, in every language you support.
Next step: prepare the reply early
The wait is a deliberate choice: in a personal conversation, a calm pause feels better than a quick reply. The next improvement we are testing is to start preparing the reply as soon as Jev says “finished”, and discard it if the user continues. That keeps the patient turn detection and hides most of the wait.
What this means if you are an IT services firm
Turn detection is one of the details that makes a voice product feel natural rather than scripted, and Jev makes the semantic part simple to build. It deserves its own design, its own settings and its own tests, beyond the transcriber’s default endpointing. The table and checklist above are a solid starting point for any voice project.
taritas builds conversational AI alongside engineering teams who are adding voice to their products. If you want a partner who has already tuned this in production, see how we work with partners.
Reading this because a client asked for voice AI? That is the conversation we are built for. What taritas does for partners.