Captioned vertical voice room
Broadcast an audio-only room as 9:16 video that reads well with the sound off.
Social audio rooms that want a vertical livestream. The hard part is not the audio — it is giving an audio-only room something to look at, with captions the viewer can actually see.
The pipeline
Sources
- livekit (audio)
- web_capture
- video_generator
Nodes
- asr
- audio_mixer
Sinks
- livekit (captions back)
- rtmp_push
livekit(room, audio) ─┬→ asr ────────────────→ livekit sink ──╮
└→ audio_mixer ──────┐ │ captions re-enter
video_generator (16x16 carrier) ───────────┴→ livekit sink │ the room as data
│
web_capture(overlay canvas) ─────┐ │
└→ rtmp_push ←──────────────╯ What makes it work
The parts that are not obvious from the JSON — usually because the naive approach does not do what you would expect.
Captions are never burned into pixels
An asr node emits sidecar data only. To make captions visible, send them somewhere that can draw: a livekit sink publishes them into the room on the avflow.asrText topic, a page joins that room and renders them, and web_capture records that page.
A caption sink still needs video
Any sink other than segment carries captions inside encoded video, so the caption sink needs a video input. A 16×16 video_generator is the cheapest legal carrier — far better than echoing the composed vertical frame back into the room.
web_capture wants a canvas
captureElement calls el.captureStream(), which exists on canvas and media elements but not on a div — so draw the scene rather than laying it out in HTML. It also lifts the frame-rate ceiling from 10 fps to 60. Omit viewport and the canvas keeps its natural size; portrait capture allows up to 1280×1920.
The Job
Replace tokens, URLs, and storage credentials, then submit it. This is the same file /examples/22-voice-room-captions.json serves.
{
"name": "voiceroom-captions",
"sources": [
{
"name": "room",
"type": "livekit",
"config": {
"serverUrl": "wss://your-project.livekit.cloud",
"token": "<subscribe-token>",
"select": {
"mediaTypes": [
"audio"
],
"excludeIdentities": [
"avflow-captions"
]
}
}
},
{
"name": "overlay",
"type": "web_capture",
"config": {
"url": "https://your-app.example.com/overlay/captions?room=voice-room",
"captureElement": "canvas#stage",
"fps": 30,
"captureAudio": false,
"waitForSelector": {
"selector": "canvas#stage[data-ready=\"true\"]",
"timeout": 20000
}
}
},
{
"name": "carrier",
"type": "video_generator",
"config": {
"width": 16,
"height": 16,
"fps": 5,
"backgroundColor": "#000000"
}
}
],
"nodes": [
{
"name": "captions",
"type": "asr",
"inputs": [
"room"
],
"config": {
"language": "multi"
}
},
{
"name": "floor",
"type": "audio_mixer",
"inputs": [
"room"
],
"config": {}
}
],
"sinks": [
{
"name": "to_room",
"type": "livekit",
"inputs": [
"carrier",
{
"name": "captions",
"select": {
"mediaTypes": [
"data"
]
}
}
],
"config": {
"serverUrl": "wss://your-project.livekit.cloud",
"token": "<publish-token>",
"videoTrackName": "avflow-caption-carrier"
}
},
{
"name": "live",
"type": "rtmp_push",
"inputs": [
{
"name": "overlay",
"select": {
"mediaTypes": [
"video"
]
}
},
"floor"
],
"config": {
"urls": [
"rtmp://ingest.example.com/live/<stream-key>"
],
"encoding": {
"videoCodec": "h264",
"audioCodec": "aac",
"videoBitrateBps": 3500000,
"audioBitrateBps": 128000,
"keyframeIntervalSec": 2
}
}
}
],
"policies": {
"maxDurationSec": 14400,
"idleTimeoutSec": 120
}
} Submit it
curl -X POST "https://api.avflow.dev/v1/jobs" \
-H "Authorization: Bearer ${AVFLOW_API_KEY}" \
-H "Content-Type: application/json" \
-d @22-voice-room-captions.json
# check status
curl "https://api.avflow.dev/v1/jobs/voiceroom-captions" \
-H "Authorization: Bearer ${AVFLOW_API_KEY}"
# stop
curl -X DELETE "https://api.avflow.dev/v1/jobs/voiceroom-captions" \
-H "Authorization: Bearer ${AVFLOW_API_KEY}"