A meeting transcript is only useful if you can answer two questions about every line: who said it, and where it sits in the recording. Speaker diarization and audio-aligned timestamps are what turn a live caption stream into something people can search, replay and quote. This guide shows how to get both over one WebSocket with the LecSync Realtime API, what the messages look like, and how the same job looks on AssemblyAI, Deepgram and OpenAI.
Everything about competitors comes from their public docs and pricing pages, read on October 6, 2026. LecSync facts come from the developers page and the API docs.
What does a meeting client need from a realtime speech API?
A meeting client needs three fields on every settled line of text: a speaker label, a start time and an end time, all measured against the same audio clock as the recording. With those, the app can colour lines by person, jump playback to any sentence, and link a summary point back to the moment it was said.
Partial text matters for the live view, but it is disposable. The record your product stores and shows later is built from final segments, so those are the ones that must carry speaker and timing.
What LecSync sends back on one socket
Turn speakers on in the connect request:
{
"transcribe": { "languages": ["auto"], "speakers": true },
"translate": { "languages": ["en", "zh"] },
"digest": true,
"store": true
}
Each settled segment then arrives as one transcript.final message with the label and the timing side by side:
{
"type": "transcript.final",
"segmentId": "gw_8f3a_2",
"speaker": "speaker_0",
"text": "So the quantization step near zero is about eight.",
"startMs": 118400,
"endMs": 122900
}
Three other messages hang off the same segmentId:
translation.finalcarries the translated text with its ownstartMsandendMs, so the translated line lands on the same row as the original.- With
aiEnhanceon, a refined translation is sent again under the samesegmentIdwithenhanced: true. Replace the row; do not append it. digest.updatepushes live minutes as sections, each withstartMs,endMsand thesegmentIdsit was built from. The docs note that these times are computed from transcript segments, not written by the model.
Partial results (transcript.partial) do not carry a speaker. LecSync only assigns a speaker once a segment settles, so render previews in a neutral style and colour them when the final arrives.
All of this costs $0.006 per audio minute. The developers page lists speaker separation, audio-aligned timestamps, translation, live structured summary and 30-day recording storage inside that one rate, and says turning features off does not lower it.
Keeping the timeline straight when connections drop
There are two kinds of reconnect in a long meeting, and they behave differently.
Upstream reconnects you never see
Recognition engines report timing relative to their own connection. If that connection is re-established mid-session, the engine's clock restarts at zero, and click-to-play starts jumping backwards. The LecSync timestamps doc explains that it measures segments on a byte clock over the audio you sent, so startMs keeps increasing through those upstream resets. With store: true, the value is a position in the downloadable recording.
The same doc is frank about one edge case: a small number of segments may arrive without timestamps when the recogniser does not supply token-level timing. Your client should accept a final with no startMs and leave it unseekable.
Client reconnects you do see
If your own socket closes, the stream has ended. The disconnects doc says there is no resumption: you call connect again, segmentId numbering and startMs both restart, and the glossary must be sent again. With store: true, each connection produces its own recording file.
The fix is a small piece of client state:
- Count the bytes you send on each stream. At 16 kHz, 16-bit mono PCM, that is 32,000 bytes per second of audio.
- When a stream closes, add its duration to a running
meetingOffsetMs. - On the new stream, store every segment as
meetingOffsetMs + startMs, and prefix its ID with the stream number. - Keep the recording files in creation order so playback can pick the right file for any meeting time.
Treat speaker labels as scoped to one stream as well. The docs do not promise that speaker_0 on a new connection is the same person, so let users confirm names after a reconnect.
How the four APIs compare for speaker-labelled live transcripts
Checked against public docs and pricing pages on October 6, 2026.
| LecSync Realtime STT | AssemblyAI Streaming | Deepgram Streaming (Nova-3) | OpenAI gpt-live-transcribe | |
|---|---|---|---|---|
| How to enable speakers | transcribe.speakers: true | speaker_labels: true | diarize_model=latest (or v1) | Not available |
| Where the label appears | On each settled segment (speaker_0…) | On each Turn (A, B…) and on each final word | On each word (integer from 0) | — |
| Timestamps | startMs/endMs in ms on final transcript and translation | Word start/end in ms from stream start | Word start/end in seconds from stream start | No word-level timestamps |
| Later label fixes | Not described in docs | SpeakerRevision at session end, optionally mid-session | Not described for streaming | — |
| Translation on the same socket | Yes, paired by segmentId | Translation add-on is pre-recorded only | No streaming translation line on pricing page | Separate gpt-realtime-translate session |
| Live summary on the same socket | Yes, digest.update with time ranges | No streaming summary add-on listed; summaries go through LLM Gateway | Summarization billed per token under Audio Intelligence | — |
| Speaker add-on price | Included | +$0.12/hr, billed on session time | $0.0020/min | — |
Sources: AssemblyAI streaming diarization, AssemblyAI pricing, Deepgram diarization, Deepgram pricing, OpenAI realtime transcription, OpenAI pricing.
Where the others are stronger
AssemblyAI goes deeper on diarization itself. Each final word carries its own speaker and a speaker_confidence score, which catches interjections inside a turn. Its SpeakerRevision message corrects earlier turns once the model has heard more of the conversation; the mid-session interval has a 120-second minimum and the docs recommend 5 minutes. Turns shorter than about one second come back as PENDING, so plan a UI state for that.
Deepgram also labels at word level, which helps if you want to split a segment where a second voice cuts in.
LecSync labels whole segments. In the default sentence display mode, the transcribe doc says segments are grouped by punctuation, speaker changes and pauses, so a change of speaker usually starts a new row. If you need word-level attribution or confidence scores, AssemblyAI or Deepgram is the closer fit.
OpenAI's guide states that gpt-live-transcribe does not return word-level timestamps, speaker labels or confidence scores, and suggests a file transcription model or an application-level fallback when you need them. It is a poor match for a speaker-labelled live meeting view.
What one hour of speaker-labelled transcription costs
Calculated from the list prices above, pay-as-you-go, one target language where offered:
| Option | Per hour | What that covers |
|---|---|---|
| LecSync | 60 × $0.006 = $0.36 | Speakers, timestamps, translation, live minutes, optional storage |
| AssemblyAI Universal-Streaming + diarization | $0.15 + $0.12 = $0.27 per session hour | Speakers and timestamps; no streaming translation |
| AssemblyAI Universal-3.6 Pro Realtime + diarization | $0.45 + $0.12 = $0.57 per session hour | Same, on the higher-end model |
| Deepgram Nova-3 Monolingual + diarization | ($0.0048 + $0.0020) × 60 = $0.408 promo; $0.582 at regular price | Speakers and timestamps |
| Deepgram Nova-3 Multilingual + diarization | ($0.0058 + $0.0020) × 60 = $0.468 promo; $0.672 at regular price | Speakers and timestamps |
| OpenAI gpt-live-transcribe | 60 × $0.017 = $1.02 | Text only, no speakers or word timestamps |
Two billing details change the real number. AssemblyAI bills streaming on how long the WebSocket is open, idle time included, and its diarization add-on applies to the whole session. LecSync bills accepted audio duration and states that no accepted audio means no charge, though its pricing doc adds that accepted silence is billed. Deepgram marks its lower streaming rates as limited-time promotions.
If speakers and timestamps are all you need, AssemblyAI Universal-Streaming is cheaper per hour. LecSync starts to win once the same meeting also needs live translation and running minutes, because those stay inside $0.36 instead of becoming a second vendor. Our 1,000-hour pricing breakdown runs that math at scale.
Building the client, step by step
- Connect.
POST https://api.lecsync.com/v1/realtime/connectwithspeakers: true, your language pair,digest: trueandstore: trueif you want click-to-play on a file. - Stream. After
session.ready, send rawpcm_s16leat 16 kHz mono, up to 64 KB per frame. No sequence numbers or offsets are needed. - Store finals by ID. Key rows on
segmentId. Filltext,speaker,startMs,endMsfromtranscript.final, then attachtranslation.finalto the same row. - Name the speakers. Show
speaker_0,speaker_1as placeholders and let the host rename them. Keep the mapping in your app, per stream. - Seek. On click, play the stored recording from
startMs.GET /v1/recordings/{sessionId}returns signed audio and transcript URLs for 30 days. - Navigate by minutes. Render each
digest.updatesection as a chapter. ItssegmentIdsand time range tell you which rows to highlight and where to jump.
A minimal handler:
ws.onmessage = (e) => {
const m = JSON.parse(e.data);
if (m.type === "transcript.final") {
rows.set(m.segmentId, {
speaker: m.speaker ?? null,
text: m.text,
startMs: m.startMs != null ? meetingOffsetMs + m.startMs : null,
endMs: m.endMs != null ? meetingOffsetMs + m.endMs : null,
});
} else if (m.type === "translation.final") {
const row = rows.get(m.segmentId);
if (row) row.translation = m.text; // enhanced: true simply overwrites
} else if (m.type === "digest.update") {
chapters = m.sections; // each push replaces the previous state
}
};
For network trouble on the live view, see how the 2-second backlog cap keeps captions from drifting behind the speaker.
Which API fits which meeting product
- Captions plus speakers, English or a single language, lowest bill: AssemblyAI Universal-Streaming with diarization.
- Word-level speaker splits and confidence for QA tooling: AssemblyAI, with Deepgram as the other word-level option.
- Speaker-labelled transcript, live translation and running minutes on one connection, one rate: LecSync.
- Spoken assistant or voice agent rather than a transcript: look at OpenAI's realtime models instead; gpt-live-transcribe alone will not give you speakers.
For broader head-to-heads, see LecSync vs AssemblyAI and OpenAI and LecSync vs Deepgram.
FAQ
Does the LecSync Realtime API support speaker diarization?
Yes. Set transcribe.speakers to true in the connect request. Every settled transcript.final then carries a speaker field such as speaker_0, alongside startMs and endMs. Partial results have no speaker because LecSync assigns it only when a segment settles. Speaker separation is included in the $0.006 per audio minute rate.
Are LecSync timestamps relative to the recording?
Yes. startMs and endMs are millisecond offsets into the stream's audio, measured from the first byte you sent, on both transcript and translation finals. They come from a byte clock, so they stay monotonic when the upstream recogniser reconnects. With store: true they map directly onto the downloadable recording.
What happens to timestamps if my WebSocket disconnects?
A closed socket ends the stream, and there is no resume. The next connect starts a new stream where segmentId and startMs restart from zero and a new recording file begins. Track the duration of each finished stream on the client and add it as an offset to keep one continuous meeting timeline.
Does speaker diarization cost extra?
On LecSync, no; it is one of the features listed inside $0.006 per audio minute. As of October 6, 2026, AssemblyAI charges +$0.12 per hour for streaming diarization on top of the model rate, and Deepgram lists streaming speaker diarization at $0.0020 per minute pay-as-you-go.
Can OpenAI's live transcription label speakers?
Not with gpt-live-transcribe. OpenAI's realtime transcription guide says the model does not return word-level timestamps, speaker labels or confidence scores, and suggests a file transcription model or an application-level fallback when those are required.
Can I give speakers real names?
The API returns labels such as speaker_0 and speaker_1. Mapping them to names is done in your app, for example by letting the meeting host rename each label. Keep that mapping per stream, because the docs do not say labels carry over to a new connection.
The short version
For a meeting client, the hard part is keeping speaker, text, translation and timing attached to the same row for an hour or more. LecSync returns all four on one segmentId, keeps the clock monotonic through upstream reconnects, and adds live minutes with time ranges, all at $0.006 per audio minute. AssemblyAI is cheaper and more detailed if you only need diarized captions. Start with the LecSync Realtime STT developers page and the speakers doc.