Does the speech core share memory and guardrails with chat?+
Yes. A call and a message about the same customer share one case and one memory, and prompt defense, fact checking and data masking run the same way on both.
What happens if a reply is slow to generate?+
A short, natural filler from DRING's approved set keeps the turn moving while the response finishes behind it, never in an apology, a number or the closing.
Can the speech to text and text to speech models be changed?+
Yes. Speech-to-text, language and text-to-speech providers are each swappable per agent and per language, and a workflow can route to a different model without a re-platform.
How does the agent know when a caller has finished talking?+
A dedicated turn-taking model handles it: barge-in lets a caller interrupt mid-sentence, and end-of-turn detection tells the agent when the caller has actually finished, so it does not talk over a pause or sit through a completed answer.
Where do the voices come from?+
From a curated library, chosen per brand and per language, with a persona set for pace, formality and tone, accent correction applied to the voice, and a pronunciation dictionary for names, places and product terms.