When to use text mode instead of the voice tester
Text mode produces a stored, addressable conversation. A live voice call
does not — there is no turn to go back to once audio has moved on. The
per-turn re-run controls only exist in text mode for that reason.
Session lifecycle: start, send, end
A text test is a server-side session. The console holds only itssessionId and the current revision.
Start binds the session to one agent.
create takes the agentId, and
nothing afterwards re-selects it — to test a different agent, end this
session and start another.
Send appends a turn. Every send returns the session’s full turn list,
its set-aside list, and a new revision. The console replaces its local
state with that response rather than appending optimistically, so what you
see is always the server’s version of the conversation.
End closes the session and clears both lists from the screen. The
screen has no procedure for re-opening a session or listing past sessions,
so treat End as final: anything you wanted to keep from a run should be
copied out before you press it. This applies to set-aside turns too.
Because the sessionId lives in component state only, reloading the
console loses the handle to a running session. Start a new one.
What a turn stores
Each turn in the session is one record:
The console renders these through the same conversation rail as the voice
tester, so a text transcript and a call transcript read identically.
How a re-run replays prior state
Pressing a turn’s re-run button callsagentTestSession.rerun with that
turnId. The session replays the conversation exactly as it stood
before that turn and then runs the turn again. The agent sees the same
history it saw the first time — not the history that accumulated
afterwards.
That is what makes the two answers comparable. If the re-run instead
appended to the end of the current conversation, the agent would answer
with extra context and you would be comparing two different situations.
Everything that came after the re-run turn is no longer part of the live
conversation, because it was produced from a reply that has now been
replaced. Those later turns are not deleted — they move to Set aside.
The response to rerun carries the new turns, the new discarded list,
and a new revision, and the console swaps all three in at once.
Editing a turn’s wording
The second button on each turn opens the wording editor. It pre-fills the turn’suserText; editing it and pressing Run again calls the same
rerun procedure with an extra text field.
Set-aside turns and comparing answers
Set-aside turns are the turns a re-run displaced. They appear in a collapsed Set aside card, rendered dimmed through the same conversation rail, and the card only shows when thediscarded list is
non-empty.
They exist so that rewinding and changing your mind does not silently
destroy work nobody chose to delete. Open the card to read the answer the
agent gave before your edit, alongside the answer it gives now.
Set-aside turns are history, not state: they are not replayed as context
for later turns, and they are not re-runnable. Ending the session clears
them.
Tool calls during a text test
Turns record the tool calls they produced. Each entry intoolCalls is a
JSON string that the console parses into three parts:
tool— the tool name shown on the rail. If the entry does not parse as JSON, the rail falls back to the generic nametooland shows the entry as supplied.args— the arguments the agent constructed. This is the field to check when you are testing whether a prompt makes the agent call the right tool with the right inputs.result— the value the tool call produced, which the turn stores alongside the arguments.
completed. The turn record carries a result
but does not distinguish how that result was produced, so do not read a
green tool-call row as proof that a downstream system was touched — verify
against that system if the side effect matters.
Tool calls belong to their turn. Re-running a turn runs it again, tool
calls included, which is worth remembering when a tool is not idempotent.
Failed turns and conversation memory
When the provider returns nothing, the turn is stored withok: false and
the rail shows an error notice instead of an assistant message: The agent
did not answer.
A failed turn is stored in the session and stays visible and re-runnable,
but it is not added to the conversation the agent remembers. The next
turn is answered as though the failed exchange never happened. This keeps a
provider hiccup from poisoning the rest of the test with a dangling user
message that never got a reply.
Press the turn’s re-run button to try it again from the same position.
The session revision
Every mutation —send, rerun, and end — sends the revision the
console last received, and every successful response returns a new one.
The revision is the session’s version number.
It guards against acting on a stale view of the conversation. A re-run is
defined relative to a specific turn list; if that list has already moved on
(a second tab sent a turn, or a re-run landed while you were typing in the
wording editor), the revision you hold no longer describes the session you
would be rewinding. The mutation is rejected rather than applied to state
you never saw, and the console surfaces the error message underneath the
composer.
The screen also disables the composer and every per-turn control while any
mutation is in flight, so the common case never produces a revision
conflict in a single tab.
Limits: what a text test does not exercise
A text test is deliberately narrow. It does not place a call and it does not carry audio, so nothing in the audio path is under test here:- No speech recognition, so you are testing the agent against clean text rather than against a transcript.
- No speech synthesis, so nothing about pronunciation, prosody, or voice choice is validated.
- No barge-in or turn-taking behaviour, and no perceived latency — everything a caller experiences as timing is absent.
- No telephony leg at all: no trunk, no codec negotiation, no call quality metrics, and no Call Detail Record.
Related
- Telemetry — where real calls emit metrics and CDRs
- Telephony Metrics — the numbers a voice test produces and a text test does not