Skip to main content

What is a Voice Step?

A Voice Step represents a conversational turn:
  1. Assistant speaks — The configured script is spoken via text-to-speech
  2. User responds — The system listens and transcribes the response
  3. AI evaluates — The LLM analyzes the response against defined conditions
  4. Navigation — The conversation moves to the appropriate next step

Voice Step Structure

All of these are set from the step’s menu. See Step-Specific Settings for the options and defaults.

Writing Scripts

The script defines what your assistant communicates. The step’s Prompt Mode (on the step menu) decides how it is used: New steps start on Instruct. We recommend Verbatim for any line you can write in advance — greetings, disclosures, read-backs and closing lines — and Instruct for turns that have to respond to what the caller just said. Older steps can show a Say badge from an earlier prompt mode. They run as Instruct steps told to say the text, so the wording can vary slightly. Switch them to Verbatim for exact wording or Instruct for guided wording.

Verbatim Scripts

Choose Verbatim when the wording must be exact: a disclosure, a legal statement, a greeting you’ve tuned carefully. Write the line exactly as it should be heard:
If the caller cuts in before a verbatim line has reached the step’s Verbatim threshold (80% unless you change it), the line is said again word for word when the step runs again; past the threshold it counts as said and the assistant carries on in its own words. See Interruptions.

Instructional Scripts

With Instruct, use describing words like “consider”, “ask about”, or “explain” for flexible delivery:
The assistant will interpret instructional scripts and generate natural responses that convey the intended meaning.

Tips for Scripts

  • Keep sentences short and clear
  • Ask one question at a time
  • Include natural transitions

Conditions

Conditions determine where the conversation goes based on the user’s response:

How Conditions Are Evaluated

The AI doesn’t do simple keyword matching. Instead, it:
  1. Understands context — Reads the caller’s reply alongside what the assistant said on this step
  2. Interprets intent — Determines what the user actually means
  3. Matches conditions — Selects the most appropriate condition
  4. Generates response — Creates a natural response incorporating the next script
This means users can express the same intent in many ways:

Special Conditions

Otherwise

The “otherwise” condition is a fallback when no specific condition matches:
  • Always include an “otherwise” condition
  • Use it to handle unexpected responses gracefully
  • On a question, point it at a step that asks again in different words; on a statement that should move straight on, point it at the next step

Ending the call from a condition

As well as a next step, a condition can end the flow or end the call. End Call lets the assistant finish speaking and still leaves the caller a short window to keep the call going; End Call (Strict) finishes the sentence and hangs up. See Where a condition can send the call.

Question Answering

If the user asks a question that can be answered from the Knowledge Base or context, the assistant can respond without leaving the current step. This is handled automatically — the assistant answers the question and then continues with the current step’s flow.

Interrupts

By default (Normal in the step’s Interruptions menu), the first few seconds of each line are protected, and after that the caller can speak over the assistant, causing it to stop and process the interruption. You can change this per step:
  • Uninterruptible: the line always plays in full. Anything said over it is answered afterwards.
  • Uninterruptible turn: nothing the caller says during the step’s turn is acted on, though it is still recorded.
  • Allow skip: if the caller speaks before the line starts, the line is dropped.
See Interruptions for when to use each. When an interrupt occurs on a Normal step, if the assistant believes it hasn’t fully conveyed its message for the step, it will attempt to repeat the key information before continuing. This ensures important details aren’t lost due to interruptions. See Assistant Behaviour for more on how the assistant handles interruptions.

Best Practices

  1. One question per step — Don’t overload the user
  2. Anticipate responses — Think about all ways users might reply
  3. Include fallbacks — Every voice step ends with an “Otherwise” (the editor won’t save one that doesn’t). On a question, point it at a step that asks again in different words; on a statement that should move straight on, point it at the next step.
  4. Keep transitions natural — Next steps should flow from the response
  5. Test with real language — Users don’t speak in keywords

Next Steps