Skip to content
Console

Data events

Captions, translations, agent turns, speaker levels, and layout state are not pixels — they travel as data events on the same inputs topology as audio and video. A node that produces them declares data in its output, and a sink that lists that node as an input forwards them.

This page is the reference for what the payloads look like and how each sink delivers them. For which node produces what, see the node pages linked below.

TypeProduced byEmitted
avflow.asrTextasrAlways
avflow.translateTexttranslateAlways
avflow.voiceAgentTextvoice_agentAlways
avflow.audioLevelsaudio_mixerOnly when audioLevels.enabled is true
avflow.videoMixerLayoutvideo_mixerOnly when layout.common.emitLayoutInfo is true

The two opt-in events cost nothing when left off — the mixer skips the work entirely rather than computing and discarding it.

{
"text": "so the migration lands on Friday",
"startMs": 12400,
"endMs": 14850,
"speaker": "Dana",
"identity": "dana-01",
"isFinal": true
}

speaker is the human-readable label (displayName, falling back to identity); identity is the routing key. Interim results arrive with isFinal: false and are superseded by a final result covering the same span — render interim text in place rather than appending it, or you will see the same sentence three times.

{
"text": "el despliegue es el viernes",
"kind": "translation",
"language": "es",
"isFinal": true
}

kind is source for the recognised original and translation for the translated text, so one stream carries both sides.

{ "text": "I can move that to Monday.", "role": "assistant" }

One final line per conversation turn. role distinguishes the agent from the participant it answered.

{
"timestamp": 1717171717000,
"speakers": [
{ "identity": "dana-01", "level": 4, "volume": 0.62 },
{ "identity": "sam-02", "level": 0, "volume": 0.01 }
]
}

level is a quantized tier (0..levels - 1, default 5 tiers) suitable for driving a UI meter; volume is the raw RMS value before quantization. Sampled every intervalMs (default 1000, range 200–2000).

{
"layoutType": "speaker",
"width": 1280,
"height": 720,
"items": [
{ "identity": "dana-01", "sourceName": "room", "x": 0, "y": 0, "w": 960, "h": 720 },
{ "identity": "sam-02", "sourceName": "room", "x": 960, "y": 0, "w": 320, "h": 180 }
],
"onKeyframe": true
}

Each item is one placed source with its pixel rect on the canvas, which is what lets a player draw name tags, click targets, or focus indicators over a flat mixed frame. Emitted when the layout changes, not per frame.

The transport differs, and it decides what your receiving code looks like.

SinkDelivery
livekitReliable publishData, topic = event type
jitsiEndpoint message, type = event type
dailyApp message, type = event type
agoraData-stream message, type = event type
rtmp_push, srt_push, whipEmbedded as text metadata in the video bitstream (H.264 SEI)
segmentavflow.asrText becomes a WebVTT rendition; other events embed as in-stream metadata
image, websocketNot supported

Two consequences worth knowing before you wire a graph:

Transports with topics send only the payload. LiveKit, Jitsi, Daily, and Agora use the event type as the channel label or message type and deliver the data object shown above. Transports without a topic concept send the whole { "type": ..., "data": ... } envelope instead, so parse accordingly.

Non-segment sinks need a video input to carry events. Metadata rides inside the video elementary stream, so a sink with no video track has nowhere to put it and the Job is rejected at submit time. segment is the exception for captions specifically, because WebVTT is a separate rendition. For an audio-only room, add a video_generator as a minimal carrier — a 16×16 track at 5 fps is enough.

Nothing extra is needed to subscribe: joining the room as a normal participant is enough, because events arrive on the standard data channel of whichever platform the sink publishes to.

// LiveKit
room.on(RoomEvent.DataReceived, (payload, _participant, _kind, topic) => {
if (topic !== 'avflow.asrText') return;
const caption = JSON.parse(new TextDecoder().decode(payload));
if (caption.isFinal) appendLine(caption.speaker, caption.text);
else showInterim(caption.speaker, caption.text);
});

Captions are not burned into the video. To show them on the pixels — for a platform that has no caption UI of its own — render them on a web page and capture that page back into the mix with web_capture. The captioned voice room use case does exactly this and links to working code.

A sink can take only the data from a node and none of its media, which is how you attach captions to a mix without also mixing the ASR node’s audio path:

{
"name": "hls",
"type": "segment",
"inputs": [
"mix_video",
"mix_audio",
{ "name": "captions", "select": { "mediaTypes": ["data"] } }
]
}

See Select filter for the full filter shape.