Turn Detection, Perception, TTS, STT, Hotwords — what they do, why we picked good defaults, and when they might matter.
Home › Help Center › AI Avatars › Behind the scenes: the settings we picked good defaults for
Building a great AI avatar normally exposes a dozen advanced settings: turn detection model, perception model, text-to-speech engine, speech-to-text engine, hotwords, tools, audio tools, visual tools. Most users don't need to touch any of them — we picked good defaults so you can focus on the prompt and persona. This article explains what each one does, in case you're curious or you've seen these terms elsewhere.
What it does: decides when the visitor has finished talking and it's the avatar's turn to reply. Bad turn detection leads to two failure modes — the avatar cuts off the visitor mid-sentence, or it waits too long after they've clearly stopped speaking.
Our default: an advanced detection model that uses both audio cues and content cues. Faster, more natural, more accurate than simple timeout-based detection.
When it might matter: if your visitors regularly pause mid-thought (technical explanations, non-native speakers thinking aloud), or if they tend to be very fast back-and-forth, you might benefit from a different model. Contact us if you need this tuned.
What it does: lets the avatar "see" and react to the visitor's camera feed. Useful when conversations involve showing physical objects, body language cues, or visual demonstrations.
Our default: the most advanced perception model, balancing visual and audio. The avatar can see, but doesn't obsess over the camera — it focuses on what's being said.
When it might matter: if your use case is purely audio (a phone-style conversation, a podcast-style interview), perception costs you nothing but does nothing for you either. Most use cases benefit from leaving it on.
What it does: turns the avatar's text replies into the voice you hear. The TTS engine is what makes the difference between a robotic monotone and a warm, expressive voice.
Our default: a high-quality TTS engine that handles natural pauses, emotional inflection, and breath sounds.
When it might matter: different TTS engines have different strengths in different languages. We've picked the engine that performs best across our supported languages overall.
What it does: turns what the visitor says into text the avatar can read. The STT engine's job is invisible until it gets a name or technical term wrong.
Our default: an auto-selecting STT engine that picks the best transcription model for the visitor's language and accent.
When it might matter: if your visitors regularly use industry jargon or proper nouns that get transcribed wrong, that's usually a Hotwords problem (see below) rather than an STT engine problem.
What it does: tells the STT engine to prioritize certain words when transcribing. Especially useful for brand names, product names, and technical terms that sound like common words.
Our default: none. We don't pre-load hotwords because what counts as a hotword is entirely brand-specific.
When it might matter: if visitors mention your brand name (or specific product names, founder names, etc.) and the avatar misunderstands, that's a hotwords use case. This isn't exposed in the persona builder yet — contact us if you need this configured.
Three reasons:
If you have a specific need for one of these to be different on your account, get in touch via the support page.
The settings on this page are real settings on the underlying platform. We're not hiding them to limit you — we're hiding them because exposing every knob would slow everyone down and confuse most users.Help Center · Learn · Privacy Policy · Terms of Service · Contact