Skip to content
TURN-TAKING DECISION API

A pause isn’t
the end of
a thought.

Give your voice agent the context to wait, check back or prepare a response. Turn live speech into clear, revisable turn-taking advice.

Explore the decisions
Streaming API Revocable adviceDeveloper preview

A ten-second turn-taking walkthrough. At two seconds, “Let me think” is followed by silence and the turn stays protected. At four seconds, the answer continues. After “That's all”, a turn-end candidate is proposed at six seconds. Speech resumes at eight seconds, revoking the candidate. Waveform peaks show speech; the signal flattens during silence. Advice remains revisable, and the application controls the response.

300k

minutes of real conversation
used to train our models

Listen to the signalSpeech, pauses and endpoint evidence

Make room for contextOptional transcript and answer cues

Keep your app in controlStructured advice, clear reasons

THE MOMENT BETWEEN TURNS

One conversation.
More than one next move.

Explore how the advice changes with a thinking pause, a complete answer or a speaker picking up their thought.

Decision walkthroughChoose a moment. Follow the advice.
Participant audio01 — 03
THE MOMENT

“Let me think about that…”

CONVERSATION TIMELINE1 / 3
SpeechPause + contextSpeech returns
POLICY STATEHOLD

Listen to the speaker

Speech observed

The speaker has the floor.

While speech is observed, the policy holds the turn open. A pause has not become a decision.

WAIT
StateHOLD
DeliveryAdvisory
Inspect the event fieldsAPI-shaped output

Each decision has a reason and a state. This walkthrough shows selected fields from the EOT advice contract; your application retains control of playback.

{
  "schema": "sharedpro.eot.advice.v1",
  "type": "eot_advice",
  "action": "WAIT",
  "event_type": "state_change",
  "state": "HOLD",
  "reason_codes": [
    "speech_observed"
  ],
  "mode": "shadow",
  "stage": 3,
  "advisory": true,
  "actionable": false
}
BUILT AROUND THE SPEAKER

The space between words
deserves some intelligence.

Combine the sound of a pause with the context around it. Keep the decision open when the evidence changes.

01 / PROTECT THE THOUGHT

Context that keeps
the turn open.

Requests for thinking time, unfinished phrases and self-corrections can protect an answer. Add caller-supplied transcript or context cues alongside the audio.

Hold protectionsAnswer-type context
02 / ADAPT THE WAIT

Patience that follows
the conversation.

Use observed continuation pauses within a session to extend waiting. Preserve time already granted to the speaker.

Session pause historyConservative extensions
03 / REVISE THE DECISION

A handover you
can reconsider.

Speech resumes. Context changes. A candidate can be revoked. Your controller receives the update and decides how to respond.

Candidate identityRevocation events

Every recommendation has a trail.

Inspect reason codes, source-aligned timing, evidence availability and candidate identity. Replay microphone or file sessions in the dedicated lab and export the decision trace.

Structured events
A FOCUSED LAYER IN YOUR VOICE STACK

Your agent. Your voice.
A clearer next move.

Connect a dedicated turn-taking service alongside your existing speech and language models. Your application owns the conversation.

Participant audioLive PCM stream
Turn-Taking APIEvidence + context + policy
Your controllerAdvice, revocations, playback
Optional transcript + context from your stack
DEVELOPER CONTRACT

Stream audio in.
Read advice as it changes.

The authenticated WebSocket endpoint accepts mono 16 kHz PCM. The Python client handles framing and event reads; your backend keeps credentials and conversation control.

Transport
WebSocket + Python client
Endpoint
/api/eot/live
Outputs
Endpoint evidence + turn advice
Start with
Shadow observation and replay

Advice stays advisory. Speech-to-text, response generation, playback and interruption handling belong to your application.

JSON · start control
{
  "type": "start",
  "schema": "sharedpro.eot.stream.v1",
  "encoding": "pcm_f32le",
  "sample_rate": 16000,
  "channels": 1,
  "mode": "shadow",
  "stage": 0,
  "answer_id": "answer-1",
  "answer_shape": "open",
  "capture_epoch": 0,
  "clock_domain": "capture_samples",
  "research_capture": false
}

Begin in observation mode. After ready, send EOT2-framed participant PCM and read events concurrently. Enable further advisory stages through your serving configuration.

WHERE TIMING MATTERS

Build for a person
on the other end.

Bring turn advice into the voice experiences you’re building.

Interviews with room to think

Account for open answers, short responses, requests for more time and a speaker correcting themselves.

Coaching that listens

Give a learner time to form an answer and keep the feedback moment in your application’s control.

Voice agents with context

Combine acoustic evidence and explicit conversation cues when deciding the next conversational move.

A FEW USEFUL DETAILS

Before you
plug it in.

Built as a focused, inspectable service for your voice stack.

What goes into a turn-taking decision?

The API combines speech activity and pause timing with supported pitch and energy changes, endpoint-model evidence, answer type and optional transcript or context cues. Evidence and policy advice are returned separately, so your application can inspect the basis of a recommendation.

Will it work with our existing voice stack?

It is a standalone authenticated WebSocket API with a Python client. Feed it participant audio alongside your existing transcription path, and optionally provide transcript revisions or context. Your speech-to-text, language model and text-to-speech services stay in your stack; this API supplies turn advice.

How does it give a speaker more time?

Explicit cues such as a request to think or an unfinished phrase protect the turn. Observed continuation pauses within the same session can extend waiting. That history does not shorten a previously granted extension or create a speaker profile across sessions.

Does a turn-end candidate make the agent speak?

A COMMIT event proposes a possible handover; it is still advisory. If speech resumes, context changes or source evidence becomes invalid, the API can revoke the candidate. Your application owns playback, interruption handling and cancellation.

What can we try today?

Use the decision walkthrough on this page to explore the behavior. The dedicated project lab supports microphone input, file replay and exportable decision traces. Contact us to discuss the developer preview and evaluate the stages relevant to your use case.

How is the API trained and evaluated?

Our models were trained on 300,000 minutes of real conversation. The integration provides shadow observation, replay and decision traces to evaluate behavior in your own setting. The current API is an experimental advisory preview; production qualification and application behavior are separate from training scale.

LET’S BUILD THE NEXT CONVERSATION

Give the next thought
room to arrive.

Explore the Turn-Taking Decision API for your voice application.

Experimental turn-taking advice. Your application retains control of agent speech.
© 2026 Sharedpro · sharedpro.co · Vadodara, India · [email protected]Making employment portable since 2020
STEP 1 OF 5 · ABOUT YOU

Let’s start with you.

A few details so we can make the conversation relevant.