Continuous Conversations
What a resuming conversation preserves, what compaction does to it, and the limits a long task can hit.
Trinity has two conversation surfaces, and the difference is memory.
| Surface | Where | What the agent remembers |
|---|---|---|
| Chat tab on Agent Detail | /agents/{name}?tab=chat | Nothing. Each message starts fresh. |
| Workspace | /workspace | Everything — tool results, mid-task state, reasoning — carried across turns. |
The Session tab is retired. It was folded into the Chat tab as a mode, and that mode has now been removed in favour of the Workspace. Old ?tab=session links redirect to /workspace?agent=<name>, and the Chat tab carries a Continue in Workspace → link, which opens the Workspace in its own browser tab. The underlying session API still exists (see For Agents); what went away is the second UI for it.
This page covers the behavior of a resuming conversation — what carries over, what compaction does to it, and the limits a long task can hit. For the Workspace UI itself, see Workspace.
Concepts
What Resuming Preserves
A resumed turn keeps the agent's working memory. A stateless turn keeps only the words. That is why long, multi-step work belongs in the Workspace: an agent that read three files in turn two still has them in turn six, instead of re-reading them.
Turns on one chat are serialized. Two simultaneous resumes of the same session could corrupt it, so a second message while one is in flight is refused rather than queued. Start another chat if you need parallel work against the same agent.
If the underlying session file has gone missing, the platform recovers automatically: it retries once as a cold turn, re-attaching the conversation history. You get an answer; the agent has the transcript but not its prior working state.
A turn is bounded by the agent's own execution timeout(default one hour, range 1 minute to 2 hours). A turn that hits it fails naming the agent's limit rather than hanging.
Spoken turns join the same memory
A voice callin a Workspace chat runs on the voice provider, not in the agent's session, so the agent's live session never heard it. Trinity closes that gap on the next typed turn: the spoken rows since the agent's last reply are prefixed to the message, so the agent knows what was said before it answers. The two cannot overlap — a call cannot start while a reply is being written, and a typed turn is refused while a call is live in that chat — so a reply never lands in the middle of a call.
Auto-compact
When Claude Code's internal history approaches roughly 85% of the model's context window, it:
Summarizes the history into a compact summary.
Replaces its in-memory history with that summary.
Continues the current turn.
The window is model-specific, not a flat 200K — a 1M-token model compacts at ~85% of 1M, a 200K model at ~85% of 200K.
What you'll notice: the turn takes a couple of minutes longer than expected, the visible message log is untouched, and working memory survives in compressed form. After several compacts in one conversation the summary loses fidelity and answers get vaguer — that's the point to start a fresh chat.
The 50-turn Agentic-Loop Cap
One turn can use up to 50 internal Claude agentic-loop iterations — read a file, edit it, run tests, retry — before failing with:
Task exceeded turn limit: Reached maximum number of turns (50).
Consider increasing max_turns_task in guardrails or breaking into smaller subtasks.This is not the number of messages in your conversation. It is the per-turn iteration budget for a single request, and a heavy twelve-step task with retries can exhaust it.
To raise it for one agent
TOKEN=$(curl -s --fail-with-body -X POST http://localhost:8000/api/token \
-d "username=admin&password=$ADMIN_PASSWORD" \
| python3 -c "import json,sys; print(json.load(sys.stdin).get('access_token') or '')")
# 403 = correct password, second factor required (no session). Use an MCP API
# key for automation — see api-reference/authentication.md.
[ -n "$TOKEN" ] || { echo "login issued no session" >&2; exit 1; }
curl -X PUT http://localhost:8000/api/agents/<agent-name>/guardrails \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"max_turns_task": 200}'Default is 50, allowed range 1–500. Higher values reduce false failures on heavy tasks and lengthen the worst-case execution time.
Clearing Working Memory
Starting a new chat gives you a fresh conversation. On an agent's Main chat, Reset archives the conversation and starts the agent cold: the archived chat stays in your list and remains resumable as an ordinary chat, and the new Main has no memory of it. Through the session API, resetclears an existing session's cached memory while keeping its visible log — the next turn is a cold turn.
Reach for a clean slate when the agent is going in circles, when you're switching topic and don't want bleed-over, or when repeated compaction has degraded its answers.
Conversations survive container restarts, and scope is per person: an agent's owner cannot read another user's conversations with it.
Known Limitations
| Limitation | Detail |
|---|---|
| Runtimes without resume fall back to replay | Codex agents have no resume primitive, so every turn replays the visible history as text. Continuity of conversation is preserved; working memory is not. |
| Restore from backup forces one cold turn | Platform backups cover the database (conversations and messages) but not the Docker volumes holding the agent's session files. After a restore, each conversation's first turn falls back to a cold turn. The visible log is preserved. |
| Long turns survive a severed connection | If the browser sleeps mid-turn, the turn keeps running server-side and the reply appears when the tab reconnects — no false failure. Very long turns may take a moment to reconcile. |
| Recovered turns lose their metrics | When a subprocess swallows the final result event, the platform recovers the reply but records no cost or duration for that turn. The answer is correct; the numbers are missing. |
| Long chats are windowed | The Workspace shows the newest turns of a very long chat (a 30-minute voice call alone is ~180 rows) and says Earlier messages in this chat aren't shown when older ones were cut. The agent's own memory is unaffected. |
For Agents
The session API is unchanged and remains available; it simply has no dedicated UI any more. Workspace conversations run on the same engine.
| Endpoint | Method | Description |
|---|---|---|
| /api/agents/{name}/session | POST | Create a new session row |
| /api/agents/{name}/sessions | GET | List sessions (caller-scoped) |
| /api/agents/{name}/sessions/{id} | GET | Get session with messages |
| /api/agents/{name}/sessions/{id}/message | POST | Send a turn (synchronous) |
| /api/agents/{name}/sessions/{id}/reset | POST | Clear the cached session so the next turn is cold (the visible log stays) |
| /api/agents/{name}/sessions/{id} | DELETE | Delete the session |
| /api/agents/{name}/guardrails | GET / PUT | Read or change max_turns_task, max_turns_chat, execution_timeout_sec |
All session endpoints return 404 when the session_tab_enabled feature flag is off. The Workspace does not consult that flag — it has its own chat surface and its own routes (see Workspace → For Agents); its Reset is POST /api/enterprise/client-portal/agents/{name}/sessions/main/reset.