← All use cases

AI co-host in the room

A speech-to-speech agent that listens, answers, and can be interrupted.

An always-available co-host that keeps a show moving when the audience is quiet, or a live assistant that answers questions out loud. One source, one node, one sink.

audio only · speech in, speech out The agent is excluded from its own source and holds a publish-only token. Drop either guard and it hears itself and never stops talking.

The pipeline

Sources

  • livekit (audio)

Nodes

  • voice_agent

Sinks

  • livekit (agent voice)
livekit(room, audio) → voice_agent → livekit sink (voice + transcript)

What makes it work

The parts that are not obvious from the JSON — usually because the naive approach does not do what you would expect.

Two guards stop the agent hearing itself

Exclude the agent identity in the source select, and give the sink a publish-only token. Drop either one and the agent transcribes its own voice and talks itself into an infinite conversation.

Barge-in is on by default

interrupt defaults to true, so a participant speaking over the agent cuts its current turn short instead of talking across it.

Instructions can be updated live

Submit the Job again under the same name to update instructions on the running agent — the same update path the co-host layout switcher uses.

The Job

Replace tokens, URLs, and storage credentials, then submit it. This is the same file /examples/24-ai-voice-agent.json serves.

{
  "name": "voice-agent-cohost",
  "sources": [
    {
      "name": "room",
      "type": "livekit",
      "config": {
        "serverUrl": "wss://your-project.livekit.cloud",
        "token": "<subscribe-token>",
        "select": {
          "mediaTypes": [
            "audio"
          ],
          "excludeIdentities": [
            "avflow-agent"
          ]
        }
      }
    }
  ],
  "nodes": [
    {
      "name": "agent",
      "type": "voice_agent",
      "inputs": [
        "room"
      ],
      "config": {
        "language": "en",
        "instructions": "You are a co-host on a live audio show. Keep replies to one or two sentences so the conversation stays quick. When the hosts talk to each other, stay quiet unless asked something.",
        "greeting": "Hey everyone, I just joined. What are we talking about?",
        "interrupt": true
      }
    }
  ],
  "sinks": [
    {
      "name": "to_room",
      "type": "livekit",
      "inputs": [
        "agent"
      ],
      "config": {
        "serverUrl": "wss://your-project.livekit.cloud",
        "token": "<publish-token>",
        "audioTrackName": "agent-voice"
      }
    }
  ],
  "policies": {
    "maxDurationSec": 7200,
    "idleTimeoutSec": 120
  }
}

Submit it

curl -X POST "https://api.avflow.dev/v1/jobs" \
  -H "Authorization: Bearer ${AVFLOW_API_KEY}" \
  -H "Content-Type: application/json" \
  -d @24-ai-voice-agent.json

# check status
curl "https://api.avflow.dev/v1/jobs/voice-agent-cohost" \
  -H "Authorization: Bearer ${AVFLOW_API_KEY}"

# stop
curl -X DELETE "https://api.avflow.dev/v1/jobs/voice-agent-cohost" \
  -H "Authorization: Bearer ${AVFLOW_API_KEY}"