Data events
Captions, translations, agent turns, speaker levels, and layout state are not
pixels — they travel as data events on the same inputs topology as audio
and video. A node that produces them declares data in its output, and a sink
that lists that node as an input forwards them.
This page is the reference for what the payloads look like and how each sink delivers them. For which node produces what, see the node pages linked below.
Event types
Section titled “Event types”| Type | Produced by | Emitted |
|---|---|---|
avflow.asrText | asr | Always |
avflow.translateText | translate | Always |
avflow.voiceAgentText | voice_agent | Always |
avflow.audioLevels | audio_mixer | Only when audioLevels.enabled is true |
avflow.videoMixerLayout | video_mixer | Only when layout.common.emitLayoutInfo is true |
The two opt-in events cost nothing when left off — the mixer skips the work entirely rather than computing and discarding it.
Payloads
Section titled “Payloads”avflow.asrText
Section titled “avflow.asrText”{ "text": "so the migration lands on Friday", "startMs": 12400, "endMs": 14850, "speaker": "Dana", "identity": "dana-01", "isFinal": true}speaker is the human-readable label (displayName, falling back to
identity); identity is the routing key. Interim results arrive with
isFinal: false and are superseded by a final result covering the same span —
render interim text in place rather than appending it, or you will see the same
sentence three times.
avflow.translateText
Section titled “avflow.translateText”{ "text": "el despliegue es el viernes", "kind": "translation", "language": "es", "isFinal": true}kind is source for the recognised original and translation for the
translated text, so one stream carries both sides.
avflow.voiceAgentText
Section titled “avflow.voiceAgentText”{ "text": "I can move that to Monday.", "role": "assistant" }One final line per conversation turn. role distinguishes the agent from the
participant it answered.
avflow.audioLevels
Section titled “avflow.audioLevels”{ "timestamp": 1717171717000, "speakers": [ { "identity": "dana-01", "level": 4, "volume": 0.62 }, { "identity": "sam-02", "level": 0, "volume": 0.01 } ]}level is a quantized tier (0..levels - 1, default 5 tiers) suitable for
driving a UI meter; volume is the raw RMS value before quantization. Sampled
every intervalMs (default 1000, range 200–2000).
avflow.videoMixerLayout
Section titled “avflow.videoMixerLayout”{ "layoutType": "speaker", "width": 1280, "height": 720, "items": [ { "identity": "dana-01", "sourceName": "room", "x": 0, "y": 0, "w": 960, "h": 720 }, { "identity": "sam-02", "sourceName": "room", "x": 960, "y": 0, "w": 320, "h": 180 } ], "onKeyframe": true}Each item is one placed source with its pixel rect on the canvas, which is what lets a player draw name tags, click targets, or focus indicators over a flat mixed frame. Emitted when the layout changes, not per frame.
How each sink delivers events
Section titled “How each sink delivers events”The transport differs, and it decides what your receiving code looks like.
| Sink | Delivery |
|---|---|
livekit | Reliable publishData, topic = event type |
jitsi | Endpoint message, type = event type |
daily | App message, type = event type |
agora | Data-stream message, type = event type |
rtmp_push, srt_push, whip | Embedded as text metadata in the video bitstream (H.264 SEI) |
segment | avflow.asrText becomes a WebVTT rendition; other events embed as in-stream metadata |
image, websocket | Not supported |
Two consequences worth knowing before you wire a graph:
Transports with topics send only the payload. LiveKit, Jitsi, Daily, and
Agora use the event type as the channel label or message type and deliver the
data object shown above. Transports without a topic concept send the whole
{ "type": ..., "data": ... } envelope instead, so parse accordingly.
Non-segment sinks need a video input to carry events. Metadata rides
inside the video elementary stream, so a sink with no video track has nowhere to
put it and the Job is rejected at submit time. segment is the exception for
captions specifically, because WebVTT is a separate rendition. For an
audio-only room, add a video_generator as a
minimal carrier — a 16×16 track at 5 fps is enough.
Receiving events
Section titled “Receiving events”Nothing extra is needed to subscribe: joining the room as a normal participant is enough, because events arrive on the standard data channel of whichever platform the sink publishes to.
// LiveKitroom.on(RoomEvent.DataReceived, (payload, _participant, _kind, topic) => { if (topic !== 'avflow.asrText') return; const caption = JSON.parse(new TextDecoder().decode(payload)); if (caption.isFinal) appendLine(caption.speaker, caption.text); else showInterim(caption.speaker, caption.text);});Captions are not burned into the video. To show them on the pixels — for a
platform that has no caption UI of its own — render them on a web page and
capture that page back into the mix with
web_capture. The
captioned voice room use case does exactly
this and links to working code.
Filtering an event edge
Section titled “Filtering an event edge”A sink can take only the data from a node and none of its media, which is how you attach captions to a mix without also mixing the ASR node’s audio path:
{ "name": "hls", "type": "segment", "inputs": [ "mix_video", "mix_audio", { "name": "captions", "select": { "mediaTypes": ["data"] } } ]}See Select filter for the full filter shape.