Skip to main content

Why Thinking Effort Matters

Language models can struggle with certain types of reasoning, particularly:
  • Numbers and calculations — Arithmetic, quantities, totals, percentages
  • Dates and scheduling — Day of week calculations, time differences, availability checks
  • Logic and comparisons — If/then reasoning, comparing options, eligibility checks
  • Multi-step reasoning — Tasks requiring several logical steps to reach a conclusion
These quantitative and logical tasks benefit significantly from deeper thinking. When the model has more time to reason, it makes fewer errors with numbers, dates, and complex logic. However, deeper thinking comes with a direct tradeoff: increased latency. Every step set to deep thinking adds delay before the assistant responds. In Voxworks, the latency differences between effort levels are modest. The step menu shows them next to each option:
  • Fast — The quickest reply, and the default for new steps
  • Normal — About half a second slower than Fast (+0.5s)
  • Deep — About a second slower than Fast (+1.0s)
These differences are small on any single step, but the cumulative effect matters if many steps use deep thinking. The key is to use deep thinking strategically — on the steps that will benefit most from accuracy.
Thinking Effort applies to the default GPT-OSS models. If a step is pinned to a model other than GPT-OSS from the step’s Models menu, that model runs with settings fixed by Voxworks and the step’s Thinking Effort has no effect. See Models.

What is Thinking Effort?

Set it from the step’s menu: Thinking Effort → Fast, Normal or Deep. The tooltip reads: “How much deep thinking the assistant will put into its response. Use in situations where the assistant is considering detailed information”. When the assistant generates a response, it can use different levels of reasoning:
  • Fast — Quick, direct responses for simple situations
  • Normal — Balanced reasoning for standard interactions
  • Deep — Deep reasoning for complex or important moments

Effort Levels

In exported script files the levels are stored as low, medium and high.

When to Use Each Level

Fast (Default)

Use for steps where you want quicker responses:
  • Acknowledgments and confirmations
  • Simple follow-up questions
  • Transitions between topics
  • Routine conversation

Normal

Use for steps that must reason about dates and times, or combine several pieces of information in one reply.

Deep

Use for a step where Normal has been tested and still gets the reasoning wrong.

Interaction with Other Settings


Best Practices

  1. Start with Fast — New steps default to Fast. Raise the effort only on steps that need it
  2. Escalate one level at a time — Move a step to Normal only when it has to reason about dates, times or several pieces of information; move to Deep only when Normal has been tested on that step and still makes mistakes
  3. Fix the prompt first — A step that answers badly usually needs a clearer script, not deeper thinking
  4. Test response quality — Verify fast effort responses are still good

Next Steps