Jev on real calls: can a decision model run our voice agents’ hangup and routing calls?
Every turn of a phone call, our voice agents make small decisions that never reach the caller’s ears: should we hang up now? and, for agents built as conversation flows, which branch did the caller just take? Today an LLM answers both, returning JSON we parse into a verdict.
Jev, from TypeSafe AI, is built for exactly this kind of job. It doesn’t write text. You give it a state and a set of typed questions, and it returns calibrated probabilities: a yes/no, or a distribution over named choices. So we asked whether it could take over these two decisions.
Short answer: with tuned thresholds, Jev agreed with our production LLMs on 95.0% of hangup checks and 88–89% of flow decisions, at a median of 380–460 ms. Out of the box it scored 92.7% and 72.2%, and the second gap turned out to be our integration, not the model. One flag needs a safety fix before Jev can make live decisions.
| Hangup check | Flow branch classifier | |
|---|---|---|
| Decisions replayed | 4,516 checks from 1,079 calls | 349 turns from 100 calls |
| Agreement at our first-cut settings | 92.7% | 72.2% |
| Agreement tuned | 95.0% | 88.5% (88–89% on held-out calls) |
| Jev median latency | 462 ms | 380 ms |
How we tested it
We didn’t route live calls through Jev. We replayed real production decisions offline.
Every call we run records the exact request each LLM decision received and the answer it gave. So for each hangup check and each flow decision, we sent Jev the same inputs production’s LLM had seen, and compared its verdict against production’s. No caller was affected.
- Hangup check: a deterministic 50% of one day’s calls, sampled by hashing the call id. That gave 4,805 calls, of which 1,079 were answered and reached at least one hangup check, for 4,516 decisions in total. Production uses gpt-4o-mini here.
- Flow classifier: 100 calls to conversation-flow agents over four days, 349 classifier turns. Production uses gpt-4.1-mini here.
- Thresholds were swept offline. Jev returns probabilities, so we stored them once and tried every cutoff without extra API calls.
One caveat applies to every number in this post: it measures agreement with our production model, not accuracy. Production is a reference, not ground truth. Where the two disagreed, we read samples by hand to see who was right.
Hangup check
Our hangup check asks six yes/no questions about the latest turn, such as did the agent say goodbye?, did the caller ask for more help? and is anything still unresolved? A fixed rule combines them into one verdict. With Jev, each question becomes a probability, and we combine those into a single probability of hanging up.
At our first-cut threshold of 0.7, Jev was extremely cautious. It never hung up where production didn’t, but it missed 63% of production’s hangups.
| Production \ Jev (threshold 0.7) | Jev: hang up | Jev: keep talking |
|---|---|---|
| Production: hang up (517) | 189 | 328 |
| Production: keep talking (3,999) | 0 | 3,999 |
Because Jev returns probabilities, moving the threshold is free:
At 0.6, agreement rises to 95.0%. A third of the missed hangups come back, at the cost of 8 false hangups out of 3,999. We chose not to go lower. A missed hangup just leaves the line open until our 15-second silence timeout. A false hangup cuts off a real person mid-conversation, and that costs far more.
What the disagreements showed
We read a random 30 of the 328 cases where production hung up and Jev didn’t. Production was right in about 24, Jev in about 2, and 4 were ambiguous. So Jev’s misses were real, and they had a single cause.
On clean goodbyes like “Goodbye!” / “Bye bye.”, Jev was sure the agent had signed off (about 0.95) and sure the caller wanted nothing more (about 0.04). But it still rated “is anything unresolved?” at 0.5–0.9, and that one flag dragged the combined probability under the threshold.
Jev’s extra hangups at lower thresholds were mostly defensible: callback wrap-ups like “I’ll call you tomorrow after six. Thank you!” / “Okay, thank you.” that production kept open because nobody said a literal goodbye. The next step is rephrasing the “unresolved” question so a scheduled callback counts as closed.
The flag that needs a safety fix
The six questions include one that asks whether the caller is on hold, meaning an automated hold message rather than a person. It never decides the hangup on its own. But on agents configured to end calls stuck on hold, two consecutive “on hold” answers do end the call.
Jev said “on hold” 247 times. Production said it 6 times. That’s the kind of disagreement that could hang up on live callers. Raising Jev’s cutoff for this one flag fixes it cleanly:
| Jev cutoff for “on hold” | Jev flags | Overlap with production’s 6 |
|---|---|---|
| 0.5 | 247 | 5 |
| 0.7 | 69 | 5 |
| 0.8 | 31 | 5 |
| 0.9 | 5 | 5 |
At 0.9, Jev’s hold flags are a near-exact subset of production’s.
Flow branch classifier
For agents built as conversation flows, each node lists conditions such as “Caller confirms they’re the right person” or “Caller says it’s a wrong number”. After every caller turn, a classifier picks the condition that was met, or says it’s unsure and stays on the node.
We gave Jev the conditions as named choices plus a “none yet” option. We also added a second yes/no question: has the agent finished this node’s task? Our LLM prompt has the same rule, because moving on before the agent has asked its question is a common failure.
Out of the box, agreement was 72.2%, and Jev moved to a branch only 54% of the times production did. That looked bad until we separated the two answers. Jev’s top choice matched production’s branch 83% of the time. Our task-complete gate, set at 0.5, was throwing most of those correct picks away.
| Setting | Agreement | Moves when production moved | Jev moves, production stayed |
|---|---|---|---|
| First cut (confidence 0.6, task gate 0.5) | 72.2% | 53.7% | 1 |
| No task gate (confidence 0.4) | 87.7% | 83.4% | 7 |
| Tuned (confidence 0.4, task gate 0.05) | 88.5% | 82.9% | 3 |
A low gate beats no gate
Removing the gate entirely lets through a specific mistake. Four of its seven false advances were the same exchange: the agent asks “Can you hear me clearly?”, the caller says yes, and Jev treats that as the answer to the node’s real question, which hasn’t been asked yet.
Jev scores task-complete at 0.01–0.02 on those exchanges, while the correct answers the 0.5 gate was blocking scored 0.1–0.47. So a gate at 0.05 keeps the good picks and still catches the premature ones.
Did we overfit?
Tuning two thresholds on 349 turns and reporting the result on the same turns would flatter it. So we split the calls in half, picked the best setting on one half, and scored it on the other, both ways round.
| Tuned on | Setting picked | Held-out turns | Held-out agreement | First-cut setting, same turns |
|---|---|---|---|---|
| Half A | confidence 0.4, task gate 0.03 | 194 | 88.1% | 71.6% |
| Half B | confidence 0.4, task gate 0.03 | 155 | 89.0% | 72.9% |
Both halves picked the same setting and held up on the calls they hadn’t seen. (0.03 and 0.05 score identically on the full sample, and we prefer the round number.)
Languages
Most of these calls were in Hindi, Tamil, or a mix of either with English. TypeSafe doesn’t document language support, so this was our biggest unknown. We saw no visible drop on Indic or code-mixed turns.
Latency
| Decision | Jev p50 | Jev p90 | Jev p95 |
|---|---|---|---|
| Hangup check | 462 ms | 934 ms | 1,194 ms |
| Flow classifier | 380 ms | 560 ms | 746 ms |
That’s roughly four times TypeSafe’s advertised ~100 ms. It was measured from India with 8 requests in flight, so it includes the round trip to TypeSafe’s servers. It still fits inside our 1.5-second budget for these decisions. The few slower calls fall back to the LLM, which keeps them safe but gives up the speed.
A bonus: the replay found a bug in our own pipeline
Reading thousands of hangup checks side by side surfaced something that had nothing to do with Jev. For agents with a knowledge base, we add retrieved reference text to the caller’s message before the agent replies. The hangup check was reading that combined message as the caller’s last words, on 7.4% of checks. Our current LLM was affected just as much, and it’s now being fixed.
Replaying production decisions through a second model turned out to be a good way to audit the first one.
What we’re doing next
- Tighten the “on hold” flag to a 0.9 cutoff before Jev makes any live decision.
- Adopt the tuned thresholds: hangup 0.6; classifier confidence 0.4 with a task gate of 0.05.
- Rephrase the “anything unresolved?” question so a callback or a goodbye counts as closed, then re-run the same 4,516 checks.
- Shadow it in production. Jev runs alongside the LLM on live calls, its verdicts are logged but not acted on, and we measure agreement and latency from our own servers for a week.
- Hand-label a sample of disagreements, to turn “agrees with our LLM” into “is right”.
What we learned
- Jev behaves like the product it claims to be. Where a job breaks down into narrow yes/no or pick-one questions, it tracked our LLMs closely, in Hindi and Tamil as well as English, in under half a second.
- Calibrated probabilities are the real feature. Every improvement in this post came from moving a threshold on numbers we already had, and we tested each move offline in seconds.
- Out-of-the-box scores can mislead. Our first classifier number, 72%, said “not ready”. The 83% top-choice match said the model was fine and our gate was wrong. Check what the model actually chose before blaming it.
- Agreement is not accuracy. Our production LLM made its own mistakes in the samples we read. A model that agrees with it 95% of the time can be better or worse than it; only labelled data will tell us which.

Leave a Reply