Overview
Nodes transform media between sources and sinks.
Available types
Section titled “Available types”| Type | Output | Description |
|---|---|---|
video_mixer | Video | Composite layout (grid, speaker, custom) |
audio_mixer | Audio | Mix multiple audio inputs |
audio_resample | Audio | Convert sample rate/channels per track (n:n) |
video_encoder | Encoded video | Encode raw video access units |
audio_encoder | Encoded audio | Encode raw PCM access units |
asr | Captions | Speech-to-text subtitles |
translate | Audio + data | Managed speech-to-speech translation |
voice_agent | Audio | AI voice agent (speech-to-speech) |
{ "name": "mix_video", "type": "video_mixer", "inputs": ["room_src"], "config": { }}An optional select on an object-form inputs entry filters which upstream tracks feed the node on that edge.
Structured output
Section titled “Structured output”asr, translate, and voice_agent always emit structured text alongside (or
instead of) media. video_mixer and audio_mixer can emit layout and speaker
level metadata when asked. See Data events for the
payloads and how each sink delivers them.
Pricing
Section titled “Pricing”Per-node rates in Node pricing.