Live caption products fail in a familiar way: the network hiccups for a few seconds, audio keeps arriving, and the transcript queue grows without a ceiling. Viewers then watch captions drift further behind the speaker—even after the link recovers—because every delayed frame still waits its turn.
LecSync’s Realtime Speech API documents a different policy. Server-side audio backlog is capped at 2 seconds. Audio beyond that cap is discarded rather than queued forever. After the disruption passes, captions return to the current sentence within about a second instead of chasing an ever-growing lag.
This article explains what a backlog cap is, why uncapped queues make latency grow, and how the 2-second rule fits the rest of LecSync’s one-connection realtime stack (latency numbers, languages, regions, and the flat $0.006 per audio minute rate).
What is a realtime caption backlog cap?
A realtime caption backlog cap is a server-side limit on how much unprocessed audio may wait during a live streaming session when the network or decoder slows down. On LecSync, that limit is 2 seconds: newer audio beyond the cap is dropped so caption delay cannot climb without bound. After recovery, the stream re-anchors on the current sentence within about one second, per the public developers page (~2026-10-04, Asia/Shanghai).
Why uncapped backlog makes latency keep growing
In an uncapped design, every audio frame that arrives during congestion is preserved. Recognition eventually processes the whole queue. That sounds complete, but for live classrooms, webinars, and meeting clients it produces a worse UX:
- Delay compounds: 0.4 s → 1.2 s → 3.1 s → 7.4 s → 15.0 s in the product’s uncapped illustration
- After the network recovers, the client still has to drain old audio before it shows what the speaker is saying now
- Product owners cannot promise a maximum caption lag under hitch conditions
A capped backlog trades completeness of historical audio for a bound on how late captions can become. For live products, that trade-off is usually the right one: missing a short stretch of audio during a stall is preferable to showing yesterday’s sentence for the rest of the talk.
Uncapped vs 2-second cap (product story)
Figures below follow LecSync’s public developers illustration for how latency behaves under hitch, rechecked ~2026-10-04.
| Behavior | Uncapped backlog | 2-second cap |
|---|---|---|
| During network slowdown | Audio queues without a published ceiling | Server backlog stops at 2 s |
| Audio beyond the limit | Kept and processed later | Discarded, not queued |
| Latency trend under hitch | Keeps accumulating (example steps up to ~15 s) | Stops growing once the cap is hit (~2.0 s in the illustration) |
| After recovery | Client may still chase old frames | Captions return to the current sentence within ~1 s |
Key insight: the cap is a flow-control choice resolved server-side. Clients keep sending; congestion handling does not require a custom client-side drain strategy beyond staying connected and pushing PCM after session.ready.
Calm-network latency (before any hitch)
The backlog rule matters most under stress. Under normal conditions LecSync publishes:
| Metric | Public figure |
|---|---|
| First word on screen after speech starts | 0.28 s |
| Extra delay to first translation token | ~88 ms |
| Languages | 60, any-to-any translation on the same connection |
| Gateway regions | 5: dal, hz, jp2, ru, uk |
| List price | $0.006 per audio minute, listed features included |
Preview text streams while recognition is still running and settles into finals when confirmed. Translation streams alongside the transcript rather than waiting for the full sentence. Language auto-detect is available so clients need not lock a source language before the session starts.
What else rides on the same WebSocket
The backlog cap is one piece of a single-connection design. On the same socket LecSync can also return:
- Speaker labels on final segments (when speakers are enabled), with timestamps
- Audio-aligned
startMsfor playback without extra clock correction - Live structured minutes via
digest: true - Glossary and domain context for product or course terms
- Optional 30-day recording/transcript storage when
store: true
There is no separate meter for enabling these surfaces: the public rate stays $0.006 per minute of accepted audio. Disabling features does not lower the minute price; enabling all of them does not raise it. Fractional seconds are charged proportionally; no accepted audio means no charge.
For a wider view of what one connection returns versus piecing vendors together, see Built-in translation vs DIY STT + MT and the one-connection comparison with AssemblyAI and OpenAI.
Mistakes teams make when designing caption lag
- Treating “never drop audio” as a live UX requirement. Archival completeness belongs on a recording path (
storeor offline upload). Live captions need a lag ceiling. - Only measuring median latency on a clean laptop Wi-Fi. Hitch behavior (VPN, campus Wi-Fi, mobile handoff) is where uncapped queues explode.
- Pushing backpressure into the browser or mobile app. LecSync documents server-side flow control; clients should keep sending within the connect contract instead of inventing unbounded local queues.
- Ignoring region choice. Specifying
dal,hz,jp2,ru, oruk(or letting routing pick a healthy node) keeps frames closer to capture and reduces how often backlog grows in the first place. - Assuming translation always adds a second hop. On LecSync, translation is paired on the same connection with ~88 ms to the first translation token—not a second vendor round-trip.
How to connect
Create a session with:
POST https://api.lecsync.com/v1/realtime/connect
Send a Bearer API key and an audio config (for example PCM s16le, 16 kHz, mono), then stream on the returned WebSocket URL and token. Full curl, JavaScript, and Python examples—including bounded queues, timeouts, and cleanup—live on the developers page.
FAQ
What does LecSync’s 2-second backlog cap do?
It limits how much unprocessed audio may wait server-side when the network slows. Beyond 2 seconds, audio is discarded instead of queued forever, so caption latency stops growing without a ceiling.
Do captions catch up after a hitch?
Yes. After the disruption passes, captions return to the current sentence within about one second, rather than draining an unbounded historical queue first.
How fast is the first caption word when the network is fine?
LecSync publishes 0.28 seconds from speech start to first word on screen, with about 88 ms added to the first translation token. Treat those as vendor product figures from the developers page.
Is the backlog cap a separate paid feature?
No. It is part of the Realtime Speech API behavior on the public developers page. Billing remains $0.006 per audio minute with listed features included.
Which regions can I pin?
Five gateways are listed: dal, hz, jp2, ru, and uk. Choose one in the handshake or let LecSync route to the nearest healthy node.
What else can I get on the same connection?
Speakers, digest minutes, glossary, audio-aligned timestamps, any-to-any translation across 60 languages, and optional 30-day storage—all on the same WebSocket meter.
Next step
If you are shipping live captions for classrooms, webinars, or meetings, design for hitch recovery as carefully as for first-word latency. LecSync’s documented 2-second backlog cap, 0.28 s first word, five regions, and flat $0.006 / audio minute rate give you a public contract for both calm and congested paths. Start at the LecSync Realtime STT docs and open a connect call against https://api.lecsync.com/v1/realtime/connect.