# Text test sessions and turn re-runs

> Test an agent's prompt in text mode: how a session stores turns, what a re-run replays, and what happens to set-aside turns.

The voice tester answers "does this sound right". The text tester answers
"does this prompt behave right". Those need different properties. The
second one needs repeatability, so a text test runs with no call and no
audio, and every turn it stores can be run again against the exact state
that turn originally saw.

That is the whole point of the screen: change one word in a system prompt,
ask the same question from the same position in the conversation, and
compare two answers instead of starting a new conversation and hoping it
drifts the same way.

## When to use text mode instead of the voice tester

| Use text mode when you are checking | Use the voice tester when you are checking |
| ----------------------------------- | ------------------------------------------ |
| Prompt wording and instruction-following | How the agent sounds |
| Whether the agent picks the right tool | Barge-in, interruption, turn-taking timing |
| The answer to the *same* question after a prompt edit | Transcription accuracy on real speech |
| Tool arguments the agent constructs | Latency the caller actually perceives |

Text mode produces a stored, addressable conversation. A live voice call
does not — there is no turn to go back to once audio has moved on. The
per-turn re-run controls only exist in text mode for that reason.

## Session lifecycle: start, send, end

A text test is a server-side session. The console holds only its
`sessionId` and the current `revision`.

| Procedure | Input | Returns |
| --------- | ----- | ------- |
| `agentTestSession.create` | `{ orgId, agentId }` | `{ sessionId, revision }` |
| `agentTestSession.send` | `{ orgId, sessionId, text, revision }` | `{ turns, discarded, revision }` |
| `agentTestSession.rerun` | `{ orgId, sessionId, turnId, revision, text? }` | `{ turns, discarded, revision }` |
| `agentTestSession.end` | `{ orgId, sessionId, revision }` | — |

**Start** binds the session to one agent. `create` takes the `agentId`, and
nothing afterwards re-selects it — to test a different agent, end this
session and start another.

**Send** appends a turn. Every send returns the session's full turn list,
its set-aside list, and a new revision. The console replaces its local
state with that response rather than appending optimistically, so what you
see is always the server's version of the conversation.

**End** closes the session and clears both lists from the screen. The
screen has no procedure for re-opening a session or listing past sessions,
so treat `End` as final: anything you wanted to keep from a run should be
copied out before you press it. This applies to set-aside turns too.

Because the `sessionId` lives in component state only, reloading the
console loses the handle to a running session. Start a new one.

## What a turn stores

Each turn in the session is one record:

| Field | Meaning |
| ----- | ------- |
| `id` | Turn id. Re-runs address a turn by this id. |
| `userText` | The text you sent. Also the label on the turn's re-run button. |
| `assistantText` | The agent's reply. Rendered only when `ok` is true. |
| `toolCalls` | Tool calls the turn produced, each a JSON string. |
| `ok` | Whether the provider returned a reply for this turn. |
| `atMs` | Turn timestamp, used to order the conversation rail. |

The console renders these through the same conversation rail as the voice
tester, so a text transcript and a call transcript read identically.

## How a re-run replays prior state

Pressing a turn's re-run button calls `agentTestSession.rerun` with that
`turnId`. The session replays the conversation exactly as it stood
*before* that turn and then runs the turn again. The agent sees the same
history it saw the first time — not the history that accumulated
afterwards.

That is what makes the two answers comparable. If the re-run instead
appended to the end of the current conversation, the agent would answer
with extra context and you would be comparing two different situations.

Everything that came after the re-run turn is no longer part of the live
conversation, because it was produced from a reply that has now been
replaced. Those later turns are not deleted — they move to **Set aside**.

The response to `rerun` carries the new `turns`, the new `discarded` list,
and a new `revision`, and the console swaps all three in at once.

## Editing a turn's wording

The second button on each turn opens the wording editor. It pre-fills the
turn's `userText`; editing it and pressing **Run again** calls the same
`rerun` procedure with an extra `text` field.

```ts
// same turn position, different question
rerun.mutate({ orgId, sessionId, turnId, revision, text: "what's my balance?" });
```

Replay semantics are identical: the edited text runs against the state
that preceded the original turn. This is the path for A/B-ing a phrasing —
ask it one way, re-run with the other wording, and compare the two answers
at the same point in the conversation. Later turns are set aside exactly
as with a plain re-run.

## Set-aside turns and comparing answers

Set-aside turns are the turns a re-run displaced. They appear in a
collapsed **Set aside** card, rendered dimmed through the same
conversation rail, and the card only shows when the `discarded` list is
non-empty.

They exist so that rewinding and changing your mind does not silently
destroy work nobody chose to delete. Open the card to read the answer the
agent gave before your edit, alongside the answer it gives now.

Set-aside turns are history, not state: they are not replayed as context
for later turns, and they are not re-runnable. Ending the session clears
them.

## Tool calls during a text test

Turns record the tool calls they produced. Each entry in `toolCalls` is a
JSON string that the console parses into three parts:

- `tool` — the tool name shown on the rail. If the entry does not parse as
  JSON, the rail falls back to the generic name `tool` and shows the entry
  as supplied.
- `args` — the arguments the agent constructed. This is the field to check
  when you are testing whether a prompt makes the agent call the right tool
  with the right inputs.
- `result` — the value the tool call produced, which the turn stores
  alongside the arguments.

The rail marks these calls `completed`. The turn record carries a result
but does not distinguish how that result was produced, so do not read a
green tool-call row as proof that a downstream system was touched — verify
against that system if the side effect matters.

Tool calls belong to their turn. Re-running a turn runs it again, tool
calls included, which is worth remembering when a tool is not idempotent.

## Failed turns and conversation memory

When the provider returns nothing, the turn is stored with `ok: false` and
the rail shows an error notice instead of an assistant message: *The agent
did not answer.*

A failed turn is stored in the session and stays visible and re-runnable,
but it is **not** added to the conversation the agent remembers. The next
turn is answered as though the failed exchange never happened. This keeps a
provider hiccup from poisoning the rest of the test with a dangling user
message that never got a reply.

Press the turn's re-run button to try it again from the same position.

## The session revision

Every mutation — `send`, `rerun`, and `end` — sends the `revision` the
console last received, and every successful response returns a new one.
The revision is the session's version number.

It guards against acting on a stale view of the conversation. A re-run is
defined relative to a specific turn list; if that list has already moved on
(a second tab sent a turn, or a re-run landed while you were typing in the
wording editor), the revision you hold no longer describes the session you
would be rewinding. The mutation is rejected rather than applied to state
you never saw, and the console surfaces the error message underneath the
composer.

The screen also disables the composer and every per-turn control while any
mutation is in flight, so the common case never produces a revision
conflict in a single tab.

## Limits: what a text test does not exercise

A text test is deliberately narrow. It does not place a call and it does
not carry audio, so nothing in the audio path is under test here:

- No speech recognition, so you are testing the agent against clean text
  rather than against a transcript.
- No speech synthesis, so nothing about pronunciation, prosody, or voice
  choice is validated.
- No barge-in or turn-taking behaviour, and no perceived latency —
  everything a caller experiences as timing is absent.
- No telephony leg at all: no trunk, no codec negotiation, no call quality
  metrics, and no Call Detail Record.

Use the voice tester, or a real call, once the prompt behaves the way you
want in text. The text tester tells you the agent decides correctly; only a
call tells you the caller has a good time. TeleQuick treats these as two
separate checks on purpose.

## Related

- [Telemetry](/platform/telemetry) — where real calls emit metrics and CDRs
- [Telephony Metrics](/glossary/metrics) — the numbers a voice test produces and a text test does not
