Skip to main content
Trinity
FAQ/Troubleshooting

Troubleshooting

Symptom → cause → fix for the most common problems. Short, grounded answers with links to the full documentation.

42 questions

Why did everything log me out all of a sudden?

The backend restarted. JWT tokens are invalidated whenever the backend restarts — after an upgrade, a docker compose restart backend, or a crash — so every active web session becomes invalid at once. Log in again through the web UI. MCP clients are affected too: if Claude Code shows the Trinity MCP server as disconnected, run /mcpin your session or restart the client to reconnect. Nothing is lost — agents, schedules, and chat history all persist across restarts. See Upgrading.

Why did signing out in one tab sign me out of my Workspace tab too?

Because they share one sign-in. The main app and every Workspace tab in the same browser use the same platform session, and open tabs follow it. Signing out in one tab ends the session in the others as well; a background tab drops it quietly rather than jumping to the sign-in page. It works the other way too: sign back in anywhere, and a Workspace tab left open from the old session picks up the new one instead of ending it. A client signed in with an email code is not sent to the operator sign-in page just because the browser also holds an expired operator login. See Workspace.

My new agent shows as unhealthy right after I created it

The most common cause for a GitHub-template agent is a failed repository clone: the container starts fine, but the template never landed in the workspace, so the health check reports an issue like "Agent identity clone failed" — a running-but-empty agent. Open the agent's Logstab and look for the git clone error near the top; it is usually an authentication failure. Verify your GitHub PAT can read the repository (private repos need a token with repo access), fix or set the per-agent PAT, then stop and start the agent — the clone re-runs on startup when the workspace has no repository yet. See GitHub PAT Setup.

My agent won't start

First check the backend logs for the actual error: docker compose logs backend --tail 100, or the agent's Logs tab if the container came up briefly. If the error names trinity-agent-network as not found, or the agent flips between stopped and restarting, someone ran docker compose down— see the next answer. Agents also claim incrementing SSH ports starting at 2222; if something else on the host occupies a port in that range, the container can fail to bind — check with lsof -i :2222. An agent that keeps dying early (out of memory against its limit, a bad env, an image regression) now surfaces as a restarting frequently health issue with an operator alert, and docker ps shows a Restarting status or an (N) restart count next to Up — read docker logs <agent-container-name> for the reason. See Monitoring — Recovery Patterns.

My agents are restart-looping or say the network is missing after docker compose down

docker compose down removes trinity-agent-network, which every agent container is attached to. A recreated network gets a new id, so the existing agent containers fail to start against it — and because agents carry the unless-stopped restart policy, Docker keeps retrying each one in a backoff loop while the roster shows them as stopped. Recreate the network, then replace the stale containers (their workspace volumes are untouched by removal):

docker compose up -d                  # recreates the network; add -f docker-compose.prod.yml on a server
docker rm -f <agent-container-name>   # repeat for each stale agent

Then start the agents again from the UI (the Start toggle), or fleet-wide with POST /api/ops/fleet/restart (the Fleet Restart action in the mobile admin at /m). In future use docker compose stop or ./scripts/deploy/stop.sh, never down. See Monitoring — Recovery Patterns.

My agents didn't come back after the host rebooted

Agents created on a current release come back on their own after a reboot or Docker restart (unless-stopped), so the ones that stayed down were created before that policy shipped and still carry Docker's default restart=no— a restart policy is creation-time, and a plain restart never changes it. Check with docker inspect -f '{{.Name}} {{.HostConfig.RestartPolicy.Name}}' <agent-container-name> and fix the whole fleet in one pass:

docker ps -a --format '{{.Names}}' | grep '^agent-' | while read -r c; do
  [ "$(docker inspect -f '{{.HostConfig.RestartPolicy.Name}}' "$c")" = "no" ] &&
    docker update --restart unless-stopped "$c"
done

docker update on a stopped container changes the policy without starting it, so a deliberately stopped agent stays stopped. Any recreate (a resource, runtime, or base-image change) also adopts the policy. Start the agents that missed the reboot from the UI, and if one comes back on a stale config or an old base image, POST /api/ops/fleet/restart recreates it properly. See Managing Agents and the restart-policy runbook.

The Workspace says the conversation is already handling a message

Only one turn can run per conversation at a time — each turn resumes the same working memory, and concurrent turns would corrupt it, so a second message while one is in flight is rejected as busy. Wait for the current turn to finish; the conversation tracks a turn-in-progress flag, so if your browser tab slept or disconnected mid-turn, the UI reattaches and shows the reply once the turn lands instead of failing. If you need parallel work against the same agent, use the Taskstab (headless executions run in parallel up to the agent's slot limit) or start a second chat. See Continuous Conversations.

Why is the Workspace microphone missing, or why does dictation fail?

The microphone appears only where dictation can work. Trinity's own transcription needs the platform's ElevenLabs key to be allowed to call speech-to-text; without that, the browser's speech engine is used, and a browser with none gets no microphone. ElevenLabs grants permissions per endpoint, so a key that speaks replies aloud can still be refused for transcription. An admin sees that under Settings → General → Voice (ElevenLabs) as cannot transcribe — reason: grant the key the Speech to Text permission at ElevenLabs, then save it again. When a transcription fails, the Workspace names the cause — missing permission, rejected key, no credits, too many voice messages, an unreadable recording, or a provider failure — and the same panel shows it as Last voice-input failurewith the provider's status, for 24 hours. You can always type instead. See Workspace → Dictation.

My scheduled task didn't run

Work down this checklist. (1) Autonomyis the master switch — if the agent's autonomy toggle is off, no schedule fires. (2) The individual schedule may be disabled — check its toggle on the Schedulestab. (3) A warning triangle inside the schedule's cron chip (tooltip Invalid cron expression) means the scheduler could not register the expression, and the schedule never fires until you fix it. (4) The agent must be running; executions against a stopped container fail. (5) If the agent has the freeze schedules if sync failing flag on and its git sync has failed 3+ times in a row, the scheduler skips firing until sync recovers. (6) If the template ships a pre-check hook, an empty result records a skippedexecution at zero cost — that's by design, and you'll see the row in the execution history. (7) Missed firings are only caught up within a 1-hour grace window after a scheduler restart; older ones are dropped. Also confirm the scheduler itself is healthy — its port is not published to the host, so probe it from inside the container: docker exec trinity-scheduler curl -sf http://localhost:8001/health. See Scheduling.

curl localhost:8001/health fails, but the scheduler seems fine

Nothing is wrong. The scheduler's port 8001 is not published to the host by any compose file, so a probe from the host can never connect. Probe it from inside its container instead: docker exec trinity-scheduler curl -sf http://localhost:8001/health (expect {"status":"healthy","active_schedules":N}). ./scripts/deploy/verify-platform.shruns the same checks and treats a failed scheduler probe as a warning for this reason. The other probes — curl -s http://localhost:8000/health for the backend and curl -s http://localhost:8080/health for the MCP server — do work from the host; a 503 from the backend with a migrations block means a pending or failed schema migration, not a dead backend. See Monitoring.

An execution has been stuck in "queued" for a long time

Queued means the agent's parallel slots are full — each agent runs at most max_parallel_tasksexecutions concurrently (default 3) and queues the rest, draining the queue as slots free up. Check what's occupying the slots on the Executionstab (running count is always live) and either wait, stop a running execution, or raise the agent's Parallel Capacity in the Settingstab. Slots have a TTL of the agent's timeout plus a 5-minute buffer, so even a wedged run releases its slot eventually, and a background cleanup pass marks stale runs failed — but it never mistakes a run that is merely waiting for an orphan: an execution admitted and parked in the backend's queue, or accepted by the agent but not yet spawned, is left alone (reported as dispatch_inflight_skipped, not failed). Queued rows that never get a slot are automatically expired to failed after 24 hours. See Executions and Agent Configuration.

My execution failed with a timeout

Every execution is bounded by the agent's execution timeout — default 3600 seconds (60 minutes), configurable per agent from 60 to 7200 seconds. Schedules can set their own timeout_seconds, but never above the agent's cap: the API rejects a schedule timeout above the agent cap with schedule_timeout_exceeds_agent_cap, and rejects lowering the agent cap below an active schedule with agent_timeout_below_active_schedules. So to give a long-running schedule more time, raise the agent's timeout first (agent Settings or PUT /api/agents/{name}/timeout), then the schedule's. Also consider whether the task should be split — loops or fan-out make long work observable in bounded pieces. See Agent Configuration.

The git sync indicator is red / sync keeps failing

Trinity polls git-enabled agents every 60 seconds by default and tracks consecutive sync failures; at 3+ failures it raises an operator alert and the sync status chip on the dashboard and agent Overview turns unhealthy. Open the agent's Git tab to see the failure and its classification — the conflict modal explains each case in plain English. Auth failures mean the agent has no working GitHub token — it is missing, invalid, or expired: set or update it (the agent's Git tab, or the platform PAT in Settings) and sync again; git picks up the new token on its next operation, without a restart. For a deadlocked history (both sides diverged with no common ancestor), use the recovery reset: it adopts the upstream main branch while preserving the workspace state the agent flagged for persistence, available via the Git conflict flow or POST /api/agents/{name}/git/reset-to-main-preserve-state. See GitHub Sync.

Git in my agent says index.lock already exists

Another git process holds the repository lock — the agent's own git, an auto-sync or a Pull/Sync in progress — or a killed one left it behind. Trinity's background status polling never takes that lock, so the poll is not the cause. Trinity also never deletes a lock inside a running container, because removing one that a live git still holds can corrupt the index. Once the same index.lockhas stayed unchanged across at least three status reads spanning 15 minutes or more, the agent's git status reports index_lock_stuck and the backend logs a warning. Restart the agent to clear it: at container start no git process can be running, so the startup script removes stale locks (including submodule and linked-worktree locks), logs each one, and the next status read records it under lock_recovery. See GitHub Sync → Sync Health Polling.

An alert says an agent's git remote still carries an embedded credential

Since the GitHub token moved out of agent remote URLs, Trinity cleans old token-bearing URLs automatically, but it never removes a token it cannot replace. This alert means no replacement token could be placed for that agent — for example on a full or read-only disk — so the URL was left as it was and the agent keeps working. Give the agent its own GitHub token on its Git tab; the next agent start finishes the cleanup. A different alert, Trinity could not check this agent's git remotes, means the cleanup could not read the workspace, which is expected for an agent running without full capabilities — check that agent's remotes by hand with the command in the operator runbook, Git remote token scrub. See Upgrading.

A push untracked files from my agent's repo

Before every push Trinity rebuilds the agent's .gitignorearound its own managed rules — defaults placed above the agent's own rules, so a !negation the agent wrote still wins, plus a protected floor for credential files and .trinity/state that cannot be overridden — and then untracks anything the rules now cover. When that changes which files are tracked, the push says so: the Git tab's toast and the commit message state the untracked and un-ignored counts and paths, the sync response carries removed_paths, unignored_paths, and shadowed_negations, and a notice (Push untracked files that now match .gitignore) lands in Operations → Needs Responseso an unattended scheduled sync can't do this for weeks unnoticed. If a path you wanted kept was untracked, add a !negation below the managed block — unless it sits beneath an excluded directory such as node_modules/, where git never descends and the negation is reported under shadowed_negationsinstead. If a newly un-ignored path was a secret, it is already in the remote's history: rotate it and remove the rule. See GitHub Sync.

My webhook suddenly returns 404 or 401

A 404 means the token in the URL is no longer valid: rotating a webhook mints a new token and the old URL 404s immediately, and revoking the webhook makes all triggers 404 until you generate a new one. Grab the current URL from the schedule's Webhook panel. A 401 means signature authentication is enabled and your request is missing or mis-computing the X-Trinity-Signature: sha256=<hex>header (HMAC-SHA256 of the raw request body with the signing secret). Note that rotating the URL also clears the signing secret — re-enable signing and update the caller with the new secret. See Webhook Triggers.

Credentials I updated aren't taking effect

Update credentials through the agent's Credentials tab (or the inject API), not by hand-editing files: the hot-reload path rewrites .env and regenerates .mcp.json on the running agent immediately, with no restart needed. Remember that .mcp.json is generated from .mcp.json.template plus .env— direct edits to .mcp.json are overwritten on the next regeneration, so change the .env value instead. For subscription tokens, hot-reload applies to the next Claude subprocess; a turn already in flight finishes on the old token. On older agent base images without the hot-reload endpoint, the platform falls back to recreating the container, which drops in-flight executions. See Credential Management.

PUT /api/settings/<key> returns 422

The generic settings route refuses to store a secret in the clear. Platform credentials — the Anthropic API key, the platform GitHub PAT, the Slack app token, client secret, and signing secret, and the Google API key — are encrypted at rest, and PUT /api/settings/{key} answers 422 for any of those keys or for any credential-shaped key (*_api_key, *_token, *_secret, *_pat, *_password, *_credentials), naming the dedicated route to use instead. Use that route, or set the value from the UI (Settings → Integrations → API Keys or → Slack Integration). Retention windows and the telemetry_sharing_*keys are refused on the same route for a different reason — each has exactly one validated write path (PUT /api/settings/ops/config from Settings → Retention, and the telemetry-sharing endpoints). A 422 from /api/settings/ops/config itself means one value failed type or range validation, which rejects the whole request. See Credential Management.

My agent's context is at 90%+ / it hit the context limit

High context usage (warning above 75%, critical above 90%) means the conversation history is close to the model's window. In a resuming Workspace conversation the agent auto-compacts on its own at roughly 85% — it summarizes the history mid-turn, which adds a couple of minutes to that turn and is normal; after several compacts, response quality degrades. When that happens, start a new chat: you get a fresh transcript, a fresh cost bucket, and an agent with clean working memory. On the stateless Chat tab, use its reset option or restart the agent container. See Continuous Conversations.

I'm getting rate-limit or auth errors from Claude mid-run

If the agent uses a shared subscription and the provider refuses a turn — a rate limit (429) or a rejected token — Trinity records the failure, switches the agent to another registered subscription, hot-reloads the token into the running container, and re-issues the turn once, so the message completes instead of failing; a subscription already known to be limited is switched before dispatch. If runs keep failing instead, check four things in Settings → Integrations → Claude Subscriptions: that more than one subscription is registered (with one, there is nothing to switch to); that Automatically switch subscriptions when usage limits are reached is on (it defaults to on); the Pressurecolumn and each row's Usageblock, which tell you which subscriptions are out — a rate-limited cell is quota, while auth events mean a rejected token that is never shown as a rate limit, so re-register it; and Fall back to the platform API key— with it on and a platform key configured, an agent with nothing to switch to is moved onto the API key rather than failing, and stays there until you reassign a subscription. Auto-switch depends on the failure surfacing as a 429 or auth error, so other failure shapes won't trigger it.

One more cause worth ruling out: a leftover ANTHROPIC_API_KEY in the agent's .env. A Claude runtime prefers an API key over the subscription token, so a stale key would authenticate every run — and every resulting failure got blamed on the subscription, marking healthy subscriptions unhealthy one after another. Trinity now strips those keys from the execution environment on subscription-backed Claude agents (GET /api/credentials/status names what is being suppressed), so a stale one is ignored rather than used. See Subscription Credentials.

A subscription's rate-limited badge won't clear

The badge (rate-limited in the Pressure column, limit on the Subscription pressure tile, sub limiton an agent's tile) means either a fresh provider reading says the subscription is refusing, or it hit a rate limit in the last 2 hours with no fresh reading to the contrary. It clears as soon as a probe says the provider is serving again: with Check subscription quota automaticallyon (the default), Trinity re-checks any subscription still wearing the badge every 5 minutes, so it never has to wait out the 2-hour window. If the badge persists, either automatic checks are off — turn them back on, or click Refreshin the row's Usage block to probe now — or the subscription really is still limited, in which case the row says when the limit returns (resets 19:10). A badge that reads auth or token invalid is not a rate limit at all: the provider rejected the token, and only re-registering it clears that. See Subscription Credentials.

My Codex agent gets 401 on every turn

The Codex CLI reads its credential from an auth.json file, not from the environment, so on an older release an API-key Codex agent could not complete a single turn even with OPENAI_API_KEYcorrectly injected — only ChatGPT-plan logins worked. Current releases log the CLI in with the injected key (OPENAI_API_KEY or CODEX_API_KEY) before the agent's first turn, and re-log in on the next turn whenever you rotate the key in .env. The fix lives in the agent base image, so upgrade Trinity, rebuild the base image (./scripts/deploy/build-base-image.sh), then stop and start the agent so it adopts the rebuilt image. If you use a ChatGPT plan instead, run codex loginyourself inside the agent — Trinity never overwrites a plan login with an API key. See Agent Runtimes — Codex authentication.

I get 403 on POST /api/token with the right password

The password was accepted, but the login is not complete: the account requires two-factor authentication, so the endpoint returns 403 with "detail": "mfa_required" and a challenge_token — and no access_token, because no session exists until the second factor is verified. (A 401 is the opposite case: credentials rejected, no challenge.) In the browser, the login page moves to the second-factor step on its own. For a script, either finish the challenge with the 2FA verify endpoint or, better, use an MCP API key (trinity_mcp_*) as the Bearer token — keys are not subject to the second-factor flow. Two-factor can start applying the moment an administrator enables the role policy, before anyone has enrolled, so an unattended credential can begin receiving 403s without its own configuration changing; two-factor itself requires an entitlement. See Authentication — Second Factor Pending.

The "Create your admin account" form never appears

Because an admin already exists. The form at /setupis shown only on an install with no admin account — a DigitalOcean 1-Click droplet created without a password, a blank ADMIN_PASSWORD brought up without the installer, or a hand-rolled backend. On a normal install start.sh requires ADMIN_PASSWORD, the backend creates the admin account from it on first start, and the form's endpoint (POST /api/setup/admin-password) refuses with 403 whenever a usable admin exists, whatever the setup flag says. Sign in as admin with the password from .env; to change it, edit that line and recreate the backend container (docker compose up -d backend, or ./scripts/deploy/start.sh --hostedon a hosted install) — every boot re-applies it, and there is no change-password form. What you see instead is the login page, then the Dashboard with the first-run setup overlay if anything is still unconfigured — its Sign-in email step binds an email to the admin. See First-Time Setup.

My hosted install shows a login page but I never set a password

Two different causes, told apart by where the install came from. If you brought a hosted compose up by hand with ADMIN_PASSWORD blank and no ADMIN_PASSWORD_SOURCE=browsermarker, the backend refuses the browser-claim path — set ADMIN_PASSWORD in .env and recreate the backend (docker compose -f docker-compose.hosted.yml up -d backend, or ./scripts/deploy/start.sh --hosted), then sign in as admin with it. If it is a DigitalOcean 1-Click droplet you have never opened— one that should have greeted you with the Create your admin accountform — a login page means someone else reached the IP first and claimed the admin: the droplet holds nothing yet, so destroy it and create another, and this time open the URL right after creation or restrict port 443 to your own IP until you have claimed it (or choose the password before first boot with cloud-init user-data or the trinity-do-create.shinstaller, which removes the window entirely). The droplet's console login banner says which path it took. See Single Server → Claim the Admin Account.

I changed ADMIN_PASSWORD in .env and restarted, but the old password still works

docker compose restart does not re-read .env— it restarts the existing container with the environment it was created with, so the backend never sees the new value. Recreate the container instead: docker compose up -d backend (add -f docker-compose.prod.yml or -f docker-compose.hosted.yml on a server, or re-run ./scripts/deploy/start.sh --hosted on a hosted install; on a droplet the file is /opt/trinity/.env). The backend adopts the value on that boot. The same applies to a droplet claimed in the browser, whose .env keeps ADMIN_PASSWORDblank on purpose — setting a value there and recreating the backend is the way to reset a forgotten password. See Single Server → Sign In and Harden.

My MCP client can't reach port 8080

On a production install that exposes only ports 80/443 — a firewalled host or a one-click cloud image — port 8080 is not reachable from outside, and it doesn't need to be: the frontend proxies /mcp to the MCP server, so point the client at https://your-domain.com/mcp instead of http://your-domain.com:8080/mcp. The URL Trinity advertises in the MCP Keys page's connection snippet and in an exposed agent's Copy connection config is auto-detected as http://<host>:8080/mcp, so on an 80/443-only install an admin should set the real one under Settings → MCP Keys → MCP Server URL (it must end in /mcp). On a local install http://localhost:8080/mcp is correct; if that fails, check the container with curl -s http://localhost:8080/health. And after a backend restart, MCP clients need to reconnect (/mcp in Claude Code). See MCP Server — Which URL to use.

Docker Desktop pegs my CPU and my fans won't stop

On Docker Desktop and other VM-based Docker runtimes (Colima, Rancher Desktop), Vector's default Docker-API log source gets stuck in a reconnect storm — the VM's log relay closes each stream immediately, Vector reconnects without backoff, and the Docker VM sits at several cores while the log file grows by gigabytes a day. Native Linux is unaffected. The fix is the file-source override: scripts/deploy/start.sh detects Docker Desktop and creates docker-compose.override.yml automatically; if yours is missing, run cp docker-compose.override.example.yml docker-compose.override.yml && ./scripts/deploy/start.sh, or force it with TRINITY_LOCAL_LOG_SOURCE=file ./scripts/deploy/start.sh. The trade-off is that aggregated logs land in a single ID-keyed local-*.json file instead of the name-split platform/agent files. See Querying Logs.

I see "database is locked" errors in the backend logs

Trinity's default database is SQLite, which allows one writer at a time — lock errors show up under heavy write load or, more often, when two backend containers are accidentally running against the same file. Check with docker ps | grep trinity-backend (expect exactly one line), then docker compose restart backend. Back up the database before major changes with scripts/deploy/backup-database.sh; a daily VACUUM (04:30 UTC) keeps the file compact. If lock errors persist under sustained load, move to PostgreSQL by setting DATABASE_URL— the migration guide is at SQLite to PostgreSQL. See Monitoring — Recovery Patterns.

Port 80 or 8000 is already in use when I start Trinity

A "port is already allocated" error on docker compose upmeans another process on the host owns that port. The frontend's host port is configurable: set FRONTEND_PORT in .env (default 80) and restart. The backend maps host port 8000 in docker-compose.yml ("8000:8000"); either stop whatever else is using 8000 or edit that mapping. Redis binds only to 127.0.0.1:6379, and agents claim incrementing SSH ports in the 2222–2262 range — if an agent fails to start, check for a squatter in that range with lsof -i :2222. See Setup.

Where do I find the logs for X?

Three places, by scope. For one agent, the Logstab on its detail page shows the container's stdout/stderr with auto-refresh. For a platform service, use Docker directly: docker compose logs -f backend (or frontend, scheduler, mcp-server). For structured, queryable history, Vector aggregates everything into daily-rotated JSON files — docker exec trinity-vector sh -c "tail -50 /data/logs/platform-$(date +%Y-%m-%d).json" | jq . for platform logs, and the matching agents-*.jsonfor agent containers (on Docker Desktop with the file-source override, it's a single local-*.json file instead). See Agent Logs and Querying Logs.

My disk keeps filling up

Check where the space went first: df -h / and docker system df. The usual suspects are unused Docker images and build layers — reclaim them with docker system prune -f (safe to run). Vector log files rotate daily and are archived automatically after the retention period (default 90 days); check sizes with docker exec trinity-vector ls -lh /data/logs/ and trigger archival early via POST /api/logs/archive (admin). The database stays lean on its own: retention sweeps prune old execution logs, execution rows, and health checks on configurable windows, and a daily VACUUM reclaims the freed pages. Treat under 20% free disk as a warning and under 5% as critical. See Monitoring and Operations Monitoring.

My retention windows changed after an upgrade

They didn't — retention is per-install configuration, and an upgrade never changes how much data an existing install keeps. On every boot Trinity writes an explicit row for any retention window that has none, at the value already in force, so upgraded installs (not just fresh ones) own their windows and a later change to the built-in defaults, in either direction, cannot silently alter them; GET /api/settings/retention reports every window with source db-row. The 5-day community floor on log windows is seeded into fresh community installs only, never applied retroactively. If a number still looks different, compare it against Settings → Retentionand the audit log, which records every retention change with the windows that moved — retention has a single write path, and Reset to defaults for the other operator settings deliberately skips it. See Monitoring — Retention Sweeps.

My loop stopped before finishing all its runs

Check the loop's stop reason — in the agent's Loops tab, or in the Workspace rail's Loops tab, where the status words say why it ended. Stopped — it stopped making progress (no_progress) means several consecutive runs returned an identical response (default threshold 3; set no_progress_threshold: 0 to disable), so Trinity assumed a doom loop rather than burn budget. Stopped — time limit reached (deadline_exceeded) means the loop's max_duration_seconds wall-clock deadline passed; it is checked between runs, so the in-flight run always finishes first. Stopped — cost budget reached (budget_exhausted) means total cost met max_cost_usd at a run boundary, and stop signal matchedmeans the agent's response contained your sentinel — until-mode working as designed. By default loops are fail-fast (on_failure: abort): the first failed iteration ends the loop as failed; set on_failure: continue to tolerate failures, bounded by max_consecutive_failures (default 3) and finishing as completed_with_errors. A backend restart no longer interrupts a loop — one between runs is picked up again on boot, and one whose run was in flight continues from that run's outcome; interrupted appears only on loops from older builds. See Agent Loops.

Usage sharing says the receiver answered 404, or that the last send went somewhere else

Both come from the Recent sends record on Settings → General → Usage sharing, which logs every attempt with the receiver it went to and the HTTP status. The receiving service answered 404 at the default address means the hosted receiver did not accept the share; the send is recorded and retried automatically on the next heartbeat (every 10–20 minutes, once a share is due), and after five consecutive failures attempts drop to at most one per half-interval until the receiver answers — a dead receiver or an air gap never affects the platform. The receiver at <receiver> answered 404. Check TELEMETRY_SHARING_URL means you pointed sharing at a non-default address that isn't accepting shares — verify the URL. The mismatch sentence (That send went to X; sharing is now configured for Y, which has not seen it) appears when the newest send went to a different address than the one configured now: changing TELEMETRY_SHARING_URL does not start a new delivery episode, so the new receiver only sees shares from the next heartbeat on. See Product Telemetry — Delivery and the receiver.

The Operations queue shows "Legacy skills-library adoption refused" and Got it doesn't clear it

The row is a platform heads-up filed by _skills-sync on an install that still carries a legacy single skills-library address matching none of its configured skill sources — adding a source is an admin action, never an automatic one, so Trinity refuses the adoption and tells you once per refused URL. Got it is the wrong control: an acknowledged row moves to Resolvedand can never be cleared from there, because it waits for a delivery to an agent that does not exist. Cancel it instead — Clear All on the Needs Response tab cancels every pending item shown, or cancel just that row with POST /api/operator-queue/{id}/cancel (the UI has no per-item Cancel button). Copies filed by earlier releases at high priority clear the same way, followed by Clear All on Resolved for any you already acknowledged. If you do want that repository, add it as a source in the skills-library panel. See Skills and Playbooks and Operations.

The canvas Share button is missing

Two conditions decide it. Sharelives on the agent's Canvastab on Agent Detail only — the Workspace rail's Canvas tab has chips, PDF, and (for an owner or admin) Manage, but sharing a canvas by link is deliberately not done from the rail. And only the agent's owner or an admincan share, pin, or delete: if you can see an agent's canvases but do not own it, there are no such controls rather than buttons that would refuse, and an external Workspace client always has a read-only panel. Open the agent's Canvas tab as its owner (or ask an admin) and the button is there. See Agent Canvas.

An agent says it can't create a new canvas

It has hit the per-agent cap: each agent can hold 100 canvases (CANVAS_MAX_PER_AGENT on the backend), and at the limit a new canvas is refused with 409and a message telling the agent to retire one first — updating any existing canvas still works, and nothing is ever deleted automatically to make room. The Canvas tab shows the count as you approach the limit. The usual cause is an agent writing a fresh canvas per run instead of keeping one per topic current: ask it to reuse a canvas, and clear the finished ones — Delete a single canvas, or switch on Manage on the Canvas tab (Agent Detail or the Workspace rail) to select several and Delete selected; deleting the default main canvas is fine, the agent recreates it on its next write. See Agent Canvas.

The usage-sharing question never came back after I skipped it

That is how a skip works. Skip — later in Settings → General on the Usage sharing step leaves the decision undecided in this browser only, and nothing reopens a closed first-run setup for a skipped step — not even the Your first scheduled run just completedre-ask, which only appears when the overlay opens for another reason. (A snooze chosen on an older release's Dashboard card counts as a skip while its 14 days last.) To decide now, use the permanent home, Settings → General → Usage sharing — the toggle, the backfill choice, and the exact payload preview — or reopen the whole sequence with Settings → General → First-run setup → Re-run setup (or add ?onboarding=1 to the Dashboard URL). If Don't ask again was chosen, the ask is silenced on every device; an admin can delete the telemetry_sharing_dismissed_at setting to bring it back. See Product Telemetry.

My agent can't reach another agent

Inter-agent calls are denied by default — an agent cannot call another until you grant permission explicitly, and grants are directional (allowing A→B does not allow B→A). The calling agent's chat_with_agenttool returns an error and the target won't even appear in its list_agents results until permitted. Open the calling agent's Permissions tab, toggle on the agent it should be allowed to call, and click Save Permissions; repeat from the other agent's tab if the conversation flows both ways. Permissions also gate shared folders and event subscriptions between agents. See Agent Permissions.

I deleted an agent by accident

Deletes are soft: the agent's container is removed and its schedules stop firing immediately, but the database records are kept and recoverable for a retention window (default 180 days) before being purged. An admin can list recoverable agents via GET /api/admin/soft-deleted/agents and restore one with POST /api/admin/soft-deleted/agents/{name}/recover. Recovery restores the records only — it does not recreate the container — so start the agent afterward with POST /api/agents/{name}/start or the UI toggle. Deleted schedules are soft-deleted the same way (default 30-day window) with matching recovery endpoints. See Managing Agents.