Customer Story

How Jack & Jill improved Jack's conversations by halving turns per call

Letting candidates finish their thought, with 50% longer utterances

September 24, 20263 min read
Reson8 x Jack & Jill

Jack & Jill's agent Jack runs candidate intake, coaching and mock interviews through voice. Candidate profiles are built from the conversation with Jack, and then passed on to Jill, the agent for hiring managers. When speaking to candidates, Jack needs to guide them through the conversation without cutting them off mid-answer and losing valuable information. Today, Jack & Jill uses Reson8's turn detection speech model to get Jack's end-of-turn detection right. The result: 50% fewer turns per call and 50% longer utterances. Candidates stay in their flow, and Jack gets richer material to build their profile on.

At a glance

CompanyJack & Jill, London. jackandjill.ai
ProductAI career agents and recruiting platform. Two agents built around one candidate profile: Jack helps candidates find a new role, and Jill acts as the AI recruiter for hiring teams, sourcing matched candidates from Jack's network.
Use caseReal-time speech-to-text and end-of-turn detection for voice agents.
DeploymentReson8 realtime and turns API via the LiveKit plugin, running on GPU infrastructure in the Netherlands. Zero audio retention.
Why Reson8Turn-taking accuracy Jack & Jill could rely on for a better candidate experience: 50% fewer turns per call and 50% longer utterances.
Scope100% of Jack & Jill's audio volume processed through Reson8: candidate calls and real-time dictation.
TimelineFirst meeting to full production rollout in 15 days.

Candidates pause mid-answer, and generic turn detection ends the turn

Jack & Jill offers an AI-driven career agent that provides candidates with a personalized experience in their search for a new job. This translates into interview coaching, salary benchmarking, and direct connections to hiring managers, with the goal of connecting the right candidates to the right roles. Jack & Jill runs two agents: candidates sign up and start a conversation with Jack, who runs them through a 15-20 minute intake call. When Jack's done collecting information, Jill matches the profile that comes out of the call to hiring managers. Today, about 40% of new users arrive through a voice call. The other 60% can use in-app dictation to quickly go through the text-based onboarding. Once onboarded, freeform conversations run anywhere from a couple of minutes to over an hour.

Turn detection in a general-purpose speech engine is built for back-and-forth conversations. Turns are short, and pauses usually mean a speaker has finished. A coaching interview often takes an entirely different shape. Candidates stop to structure their thinking, rethink their answer as they're speaking, and then carry on to fully form their thought. A speech model that reads the pause as the end of the turn interrupts the candidate's train of thought and gives Jack much less information to build the conversation and the candidate profile upon.

Jack & Jill runs a cascaded voice pipeline, meaning speech-to-text output is fed into an LLM and turned into text-to-speech, rather than a speech-to-speech model which does not produce a transcript.

“For Jack, both the words and the turn detection have to be right. In a coaching interview, the whole value is in letting someone talk freely without being interrupted, to keep them in their flow. Reson8's end-of-turn detection really stood out for just that.”Ainhoa Arias, Founding Operator, Jack & Jill

Jack & Jill deployed the Reson8 model to 100% of volume 15 days after the first meeting

Jack & Jill measured results on its own calls. Turns detected per call dropped by 50%, while individual utterances ran 50% longer. Candidates stayed in their answers instead of being cut off halfway through. Call duration held roughly steady, so candidates shared more without spending longer on the phone, and the profiles Jill works from came out meaningfully richer. Word error rate came in well below the previous engine too. Jack & Jill moved from an A/B test to 100% of its volume within 15 days.

The Reson8 team suggested taking it one step further and embedding dynamic turn-taking in the pipeline. Jack now adapts to the rhythm of each candidate rather than running one fixed setting across every call. Someone who thinks out loud in pauses gets a slower pace than someone who answers in quick bursts and drives a faster conversation. This is true of both speech-to-text and text-to-speech, so conversations feel truly natural.

“Reson8's model lets us capture a richer context than what we previously could. Candidates now speak uninterrupted for much, much longer. They go down a rabbit hole and tell us things they'd never have got to otherwise. These insights help better guide the conversation and build a more complete profile for them.”Ainhoa Arias, Founding Operator, Jack & Jill

Since going live, Jack & Jill has added a second use case to its product with real-time dictation: candidates can now see what they're saying transcribed in real time.

Running a voice agent that ends turns before the speaker's finished talking? Run your call audio through Reson8 and compare turns per call and word error rate against your current engine. Get in touch or get started at console.reson8.dev.

About Reson8

Reson8 gets the words right that matter most. Generic speech models mishear medication names, order numbers, addresses, and case references. Reson8’s speech-to-text adapts to your domain-specific terminology in real time or from a single text upload, no fine-tuning or audio samples needed.

Teams use it for voice agents, medical scribes, meeting notetakers, contact center analysis, captioning, legal transcription, and more. Built for European languages including English, running on our own GPU infrastructure in the Netherlands: zero audio retention, no training on customer data.

Reson8 is the speech-to-text (STT) product of Resonate Labs B.V., Amsterdam, backed by Balderton Capital and NP-Hard Ventures. Not affiliated with other companies named Reson8.