Skip to main content
Trinity
Guides/Continuous Conversations

Continuous Conversations

What a resuming conversation preserves, what compaction does to it, and the limits a long task can hit.

Trinity has two conversation surfaces, and the difference is memory.

SurfaceWhereWhat the agent remembers
Chat tab on Agent Detail/agents/{name}?tab=chatNothing. Each message starts fresh.
Workspace/workspaceEverything — tool results, mid-task state, reasoning — carried across turns.

The Session tab is retired. It was folded into the Chat tab as a mode, and that mode has now been removed in favour of the Workspace. Old ?tab=session links redirect to /workspace?agent=<name>, and the Chat tab carries a Continue in Workspace → link, which opens the Workspace in its own browser tab. The underlying session API still exists (see For Agents); what went away is the second UI for it.

This page covers the behavior of a resuming conversation — what carries over, what compaction does to it, and the limits a long task can hit. For the Workspace UI itself, see Workspace.

Concepts

•Resume — Each turn reattaches to the same underlying agent session rather than replaying the transcript as text. That preserves strictly more than what was said: files the agent read, commands it ran, where it was in a multi-step skill, and its reasoning state.
•Cold turn — A turn with no session to resume: the first message in a chat, the first turn after working memory is cleared, or a recovery turn after the underlying session file is gone. A cold turn re-sends the visible history as text so the agent still has the thread of the conversation.
•Auto-compact — Claude Code's own mid-turn summarization of its history when it approaches the model's context limit.

What Resuming Preserves

A resumed turn keeps the agent's working memory. A stateless turn keeps only the words. That is why long, multi-step work belongs in the Workspace: an agent that read three files in turn two still has them in turn six, instead of re-reading them.

Turns on one chat are serialized. Two simultaneous resumes of the same session could corrupt it, so a second message while one is in flight is refused rather than queued. Start another chat if you need parallel work against the same agent.

If the underlying session file has gone missing, the platform recovers automatically: it retries once as a cold turn, re-attaching the conversation history. You get an answer; the agent has the transcript but not its prior working state.

A turn is bounded by the agent's own execution timeout(default one hour, range 1 minute to 2 hours). A turn that hits it fails naming the agent's limit rather than hanging.

Spoken turns join the same memory

A voice callin a Workspace chat runs on the voice provider, not in the agent's session, so the agent's live session never heard it. Trinity closes that gap on the next typed turn: the spoken rows since the agent's last reply are prefixed to the message, so the agent knows what was said before it answers. The two cannot overlap — a call cannot start while a reply is being written, and a typed turn is refused while a call is live in that chat — so a reply never lands in the middle of a call.

Auto-compact

When Claude Code's internal history approaches roughly 85% of the model's context window, it:

1

Summarizes the history into a compact summary.

2

Replaces its in-memory history with that summary.

3

Continues the current turn.

The window is model-specific, not a flat 200K — a 1M-token model compacts at ~85% of 1M, a 200K model at ~85% of 200K.

What you'll notice: the turn takes a couple of minutes longer than expected, the visible message log is untouched, and working memory survives in compressed form. After several compacts in one conversation the summary loses fidelity and answers get vaguer — that's the point to start a fresh chat.

The 50-turn Agentic-Loop Cap

One turn can use up to 50 internal Claude agentic-loop iterations — read a file, edit it, run tests, retry — before failing with:

Task exceeded turn limit: Reached maximum number of turns (50).
Consider increasing max_turns_task in guardrails or breaking into smaller subtasks.

This is not the number of messages in your conversation. It is the per-turn iteration budget for a single request, and a heavy twelve-step task with retries can exhaust it.

To raise it for one agent

TOKEN=$(curl -s --fail-with-body -X POST http://localhost:8000/api/token \
  -d "username=admin&password=$ADMIN_PASSWORD" \
  | python3 -c "import json,sys; print(json.load(sys.stdin).get('access_token') or '')")
# 403 = correct password, second factor required (no session). Use an MCP API
# key for automation — see api-reference/authentication.md.
[ -n "$TOKEN" ] || { echo "login issued no session" >&2; exit 1; }

curl -X PUT http://localhost:8000/api/agents/<agent-name>/guardrails \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"max_turns_task": 200}'

Default is 50, allowed range 1–500. Higher values reduce false failures on heavy tasks and lengthen the worst-case execution time.

Clearing Working Memory

Starting a new chat gives you a fresh conversation. On an agent's Main chat, Reset archives the conversation and starts the agent cold: the archived chat stays in your list and remains resumable as an ordinary chat, and the new Main has no memory of it. Through the session API, resetclears an existing session's cached memory while keeping its visible log — the next turn is a cold turn.

Reach for a clean slate when the agent is going in circles, when you're switching topic and don't want bleed-over, or when repeated compaction has degraded its answers.

Conversations survive container restarts, and scope is per person: an agent's owner cannot read another user's conversations with it.

Known Limitations

LimitationDetail
Runtimes without resume fall back to replayCodex agents have no resume primitive, so every turn replays the visible history as text. Continuity of conversation is preserved; working memory is not.
Restore from backup forces one cold turnPlatform backups cover the database (conversations and messages) but not the Docker volumes holding the agent's session files. After a restore, each conversation's first turn falls back to a cold turn. The visible log is preserved.
Long turns survive a severed connectionIf the browser sleeps mid-turn, the turn keeps running server-side and the reply appears when the tab reconnects — no false failure. Very long turns may take a moment to reconcile.
Recovered turns lose their metricsWhen a subprocess swallows the final result event, the platform recovers the reply but records no cost or duration for that turn. The answer is correct; the numbers are missing.
Long chats are windowedThe Workspace shows the newest turns of a very long chat (a 30-minute voice call alone is ~180 rows) and says Earlier messages in this chat aren't shown when older ones were cut. The agent's own memory is unaffected.

For Agents

The session API is unchanged and remains available; it simply has no dedicated UI any more. Workspace conversations run on the same engine.

EndpointMethodDescription
/api/agents/{name}/sessionPOSTCreate a new session row
/api/agents/{name}/sessionsGETList sessions (caller-scoped)
/api/agents/{name}/sessions/{id}GETGet session with messages
/api/agents/{name}/sessions/{id}/messagePOSTSend a turn (synchronous)
/api/agents/{name}/sessions/{id}/resetPOSTClear the cached session so the next turn is cold (the visible log stays)
/api/agents/{name}/sessions/{id}DELETEDelete the session
/api/agents/{name}/guardrailsGET / PUTRead or change max_turns_task, max_turns_chat, execution_timeout_sec

All session endpoints return 404 when the session_tab_enabled feature flag is off. The Workspace does not consult that flag — it has its own chat surface and its own routes (see Workspace → For Agents); its Reset is POST /api/enterprise/client-portal/agents/{name}/sessions/main/reset.

See Also

•Workspace — the UI where continuous conversations live
•Agent Chat — the stateless Chat tab on Agent Detail
•Agent Runtimes — which runtimes support resume