Skip to main content
Trinity
Guides/Advanced Features

Advanced Features

Voice chat, outbound VoIP telephony, image generation, agent avatars, and agent-defined dynamic dashboards.

Voice Chat

Real-time voice conversations with agents, inside a Workspace chat. Audio streams bidirectionally through a backend WebSocket proxy to the Gemini Live API (~280ms latency). Gemini handles speech-to-speech; the agent itself remains the reasoning engine and is invoked on demand via tool calling — as the agent, in the chat the call belongs to.

1

Start the call one of two ways: Talk in the Agent Detail header, or the call button — the leftmost control in the Workspace composer — in whichever chat the call should belong to.

2

The orb takes the conversation column and the agent's canvas takes the right column; the header, tabs and composer stay visible but inert.

3

Speak — audio is captured as PCM 16kHz and streamed to the backend WebSocket.

4

The backend proxies audio to the Google Gemini Live API.

5

Agent response audio (PCM 24kHz) plays back in real-time.

6

When real work is needed, Gemini calls run_task — it runs as the agent, in that chat, with its skills, files and memory, and the reply lands there as a turn.

7

End the call and the spoken turns are already in the chat, as one collapsed Voice call · N min block — the agent's next typed turn knows what was said.

Requirement: a Gemini key in Settings → Integrations (Platform Keys) or GEMINI_API_KEY in .env, plus a signed-in platform user. External clients signed in with an email code do not get a call button.

Configuration

VariableDescription
VOICE_ENABLEDEnable or disable voice chat
VOICE_MODELGemini model to use for voice
WORKSPACE_VOICE_MAX_DURATIONMaximum Workspace call duration in seconds (default 1800 — 30 minutes)

Voice API

EndpointMethodDescription
/api/enterprise/client-portal/agents/{name}/voice/startPOSTStart a Workspace voice call bound to a chat (platform users only)
/api/agents/{name}/voice/startPOSTThe original per-agent start route, retained for API clients
/api/agents/{name}/voice/stopPOSTStop a voice session
/api/agents/{name}/voice/statusGETGet session status
/ws/voice/{session_id}WebSocketBidirectional audio bridge

While you talk, the agent draws on its own canvas beside the orb — diagrams, images, formatted text — and what it drew stays on the rail's Canvas tab after the call. A call lasts at most 30 minutes, a room (a chat with several agents) has no voice mode, and the per-agent voice-and-canvas page at /agents/{name}/workspace is retired. See the full Voice Chat guide.

Outbound Phone Calls (VoIP)

Agents can place real outbound phone calls. The agent dials a number through Twilio and holds a live, interruptible spoken conversation powered by Gemini Live; after you hang up, the transcript flows back to the agent so it can act on what was discussed.

This is distinct from Voice Chat: voice chat is you talking to your agent in the browser; VoIP is the agent calling a phone number over the public telephone network. This release is outbound only— agents place calls; they do not answer incoming ones.

•Bring your own Twilio — each agent owner configures a per-agent Twilio voice binding; calls bill to that account.
•Flag-gated — requires VOIP_ENABLED=true and a GEMINI_API_KEY; off by default.
•MCP tool — agents trigger calls via call_user; rate-limited and daily-capped.

Full setup, API reference, and limitations in the VoIP Telephony guide.

Image Generation

Platform image generation via a two-step Gemini pipeline: prompt refinement then image generation.

1

Submit an image generation request via API.

2

Prompt Refinement — Gemini refines the user's prompt using best-practice templates for the use case.

3

Image Generation — Gemini generates the image from the refined prompt. Returned as base64 or URL.

Used internally for agent avatars and other platform features. API: POST /api/images/generate

Agent Avatars

AI-generated avatars for agents using reference images, emotion variants, and default generation.

•Reference Image — Upload a reference image and the avatar is generated in that style.
•Variation Regeneration — Generate new variations from an existing avatar.
•Emotion Variants — The Agent Detail page cycles through emotion-based avatar variants every 30 seconds.
•Default Avatar Generation — The Generate Default Avatars button in Settings (admin) generates robot/android-style avatars for all agents without a custom avatar.
•WebP Conversion — Avatars are converted to WebP via Pillow for optimization.

API: GET /api/agents/{name}/avatar (serve) and POST /api/agents/{name}/avatar (generate/upload).

Dynamic Dashboards

Agent-defined dashboards via dashboard.yaml with 11 widget types, historical tracking, and sparkline charts.

Widget Types

11 supported types: metric, status, progress, text, markdown, table, list, link, image, divider, spacer. That is the closed set — Trinity has never had a chart, badge, or countdown widget. Trend lines come from the platform: give a metric or progress widget a stable id and its history is drawn as a sparkline automatically.

How It Works

1

The agent writes a dashboard.yaml file to its workspace.

2

The file defines widgets with type, title, value, and optional configuration.

3

Open the agent detail page and select the Dashboard tab to see the widgets.

4

Auto-refresh updates values as the agent modifies the YAML file.

5

Historical values are tracked automatically — sparklines appear for metrics with enough data points. Trend indicators show up/down/stable arrows with percentage change.

A Platform Metrics section appears at the bottom of every dashboard, auto-injected with Tasks 24h, Success Rate, Cost, and Health. This section is not controlled by the YAML file (set platform_metrics: false at the top level to opt out).

Agents control their dashboard entirely by writing to dashboard.yaml. No API call is needed — the file is read on each dashboard request. API: GET /api/agent-dashboard/{name}.