Most realtime speech-to-text price lists lead with one number: what it costs to turn live audio into text. Production apps rarely stop there. A meeting, webinar or classroom client usually also needs live translation, a speaker label on every line, and some form of running summary. Each of those can be its own meter, sometimes from a second vendor, and the headline per-minute rate stops predicting the monthly bill.
This guide prices one concrete workload against six public rate cards, shows every step of the arithmetic, and says plainly where LecSync is not the cheapest option. All figures were read from each vendor's public pricing or documentation page on October 5, 2026. Promotional rates move, so re-check the linked pages before you sign anything.
What does flat per-minute pricing mean for a realtime STT API?
Flat per-minute pricing means one rate per minute of audio covers every feature on the stream, so switching on translation, speaker labels or live minutes adds nothing to the invoice. LecSync Realtime STT lists $0.006 per audio minute with transcription, translation, speaker separation, glossary, live structured summary and 30-day storage included, billed on accepted audio duration.
Add-on pricing is the opposite model: a base speech-to-text rate, then a separate line for each extra capability. It is cheaper when you need only the base, and it grows quickly once a product spec asks for three or four features at once.
The workload we priced
To keep the comparison honest, every vendor gets the same brief:
- 1,000 hours of live audio per month, which is 60,000 minutes
- Live captions in the spoken language
- Live translation into one target language
- A speaker label on each final line
- A running summary or minutes (during the session where a vendor offers it, after the session where that is the only option)
- Pay-as-you-go list prices, no committed-use or enterprise discounts
Several vendors do not translate on the live stream. For those, we add Google Cloud Translation NMT as a stand-in second meter at its public rate of $20 per million characters after the first 500,000 free each month. That needs a character count, so here is our assumption, stated once: conversational speech runs around 150 words a minute, roughly 50,000 characters per hour. For 1,000 hours that is about 50 million characters, or 49.5 million billable after the free allowance. Dense lectures will run higher; meetings with long pauses lower.
Public prices, checked October 5, 2026
| Vendor and line item | Public list price | Billing unit |
|---|---|---|
| LecSync Realtime STT, all features | $0.006 | per accepted audio minute |
| Deepgram Nova-3 Monolingual, streaming | $0.0048 promo / $0.0077 regular | per minute |
| Deepgram Nova-3 Multilingual, streaming | $0.0058 promo / $0.0092 regular | per minute |
| Deepgram speaker diarization, streaming | +$0.0020 | per minute |
| Deepgram keyterm prompting, streaming | +$0.0013 | per minute |
| AssemblyAI Universal-Streaming (English or Multilingual) | $0.15 | per hour of session time |
| AssemblyAI Universal-3.6 Pro Realtime | $0.45 | per hour of session time |
| AssemblyAI streaming diarization | +$0.12 | per hour |
| OpenAI gpt-live-transcribe | about $0.017 | per minute (estimated cost on pricing page) |
| OpenAI gpt-realtime-translate | about $0.034 | per minute |
| Soniox real-time | $2 / 1M input audio tokens, $4 / 1M output text tokens (Soniox says about $0.12 per hour) | tokens |
| Google Cloud Speech-to-Text v2, Standard | $0.016 | per minute, first 500,000 minutes a month |
| Google Cloud Translation, NMT | $20 | per 1M characters after 500,000 free |
Deepgram marks its lower streaming rates as limited-time promotional pricing. We show both columns because the regular rate is what a long-term budget should assume.
Worked cost for 1,000 hours a month
LecSync
60,000 minutes × $0.006 = $360 a month. Translation, speaker labels, glossary and the live summary (sections of minutes pushed over the same WebSocket while the session runs) are in that number. The developers page also states that turning features off does not lower the price and turning all of them on does not raise it.
Soniox
Soniox bills tokens, so we use the usage reference on its own pricing page: one hour of audio is about 30,000 input audio tokens, and one hour of speech is about 15,000 output text tokens.
- Audio in: 30,000 × $2 / 1M = $0.06 per hour
- Transcript out: 15,000 × $4 / 1M = $0.06 per hour
- Translation into one language: roughly another 15,000 output tokens, so about +$0.06 per hour (our estimate)
That comes to roughly $0.18 per hour, or about $180 a month. Soniox says diarization and language identification are bundled and translation runs in the same real-time call. A live summary is not on its pricing page, so budget a separate LLM call.
AssemblyAI
- Universal-Streaming Multilingual plus streaming diarization: $0.15 + $0.12 = $0.27 per hour × 1,000 = $270
- Live translation: the Translation add-on (+$0.06/hr) is listed for pre-recorded audio only, so the live path needs a second service: 49.5M characters × $20 / 1M = about $990
- Summaries: the older summarization add-on is deprecated, and AssemblyAI points you to its LLM Gateway, billed per token
Subtotal: about $1,260 a month plus LLM tokens. On Universal-3.6 Pro Realtime the speech line becomes ($0.45 + $0.12) × 1,000 = $570, for about $1,560 before tokens. AssemblyAI also bills streaming on WebSocket session time, idle time included, and applies streaming add-ons to the whole session. If your sockets stay open 1,100 hours to carry 1,000 hours of speech, you pay for 1,100.
Deepgram
- Nova-3 Multilingual, promo: 60,000 × $0.0058 = $348
- Streaming diarization: 60,000 × $0.0020 = $120
- Speech subtotal: $468 at promo, or $552 + $120 = $672 at the regular rate
- Keyterm prompting, the closest match to a glossary: +$78 (60,000 × $0.0013)
- Live translation: no streaming translation line on the pricing page, so the same second meter applies, about $990
- Summaries: listed under Audio Intelligence at $0.0003 per 1K input tokens and $0.0006 per 1K output tokens, as a separate request
Subtotal: about $1,458 a month at promo rates, about $1,662 at regular rates, before keyterms and summary tokens.
OpenAI
The translation session is the fit here. It streams the source transcript, the translated transcript and translated audio in one session: 60,000 × $0.034 = $2,040 a month for one target language. OpenAI's guide recommends one translation session per output language, so a second language doubles that line. If you only want captions, gpt-live-transcribe is 60,000 × $0.017 = $1,020, and its guide says it does not return speaker labels or word-level timestamps. Summaries would come from a text model at token rates.
Google Cloud
Speech-to-Text v2 Standard: 60,000 × $0.016 = $960. Add NMT translation at about $990 for about $1,950 a month. The Speech-to-Text pricing page does not list separate diarization or summary lines, so check the docs for what your chosen model returns.
Side by side
| Stack (1,000 h/month) | Speech + speakers | Live translation | Summary | Estimated month |
|---|---|---|---|---|
| LecSync | included | included | included, live on the same socket | $360 |
| Soniox | about $120 | about $60 (token estimate) | separate LLM | about $180 + LLM |
| AssemblyAI Universal-Streaming + Google NMT | $270 | about $990 | LLM Gateway tokens | about $1,260 + tokens |
| Deepgram Nova-3 Multilingual (promo) + Google NMT | $468 | about $990 | token-billed request | about $1,458 + tokens |
| OpenAI gpt-realtime-translate | source transcript included; no diarization line | $2,040 | text-model tokens | about $2,040 + tokens |
| Google STT v2 + NMT | $960 (speakers not priced) | about $990 | not listed | about $1,950 |
The translation line moves most. Pick a cheaper MT engine and the AssemblyAI and Deepgram rows drop. What stays is the structure: two vendors, two invoices, two retry paths, and caption timing you align yourself. We walked through that trade-off in built-in translation vs a DIY STT + MT pipeline.
Where the headline rate misleads
Session time versus audio time
AssemblyAI bills the time a socket is open. LecSync bills accepted audio duration, charges fractional seconds proportionally, and charges nothing when no audio is accepted. If your clients hold connections open between turns, the two models diverge.
Each feature adds its own rate
Deepgram diarization alone adds about 35% to the promo Nova-3 Multilingual rate ($0.0020 on $0.0058). Add keyterms and the per-minute figure is already well above the number on the landing page.
Languages multiply too
OpenAI's translation guide calls for one session per output language. LecSync configures each stream with one language pair, so a second target language also means a second stream at the same $0.006 rate. Soniox's token bill grows with every extra line of output text. Count target languages before comparing.
Token billing tracks talk density
A fast-talking seminar produces more output tokens per hour than a slow support call. Per-minute billing stays the same either way, which makes forecasting easier for finance.
Promotions end
At this volume, Deepgram's regular Nova-3 Multilingual rate adds $204 a month over the promo.
When each option is the better deal
- Raw captions, one language, no translation. AssemblyAI Universal-Streaming at $0.15 an hour ($150 a month here) and Deepgram Nova-3 Monolingual at the $0.0048 promo ($288) both cost less than LecSync's $360. If that is the whole spec, choose one of them. Our Deepgram comparison covers the details.
- Transcription, translation and speakers at the lowest list price. Soniox's token pricing works out around $180 for this workload, half of LecSync. If you do not need live minutes, glossary, translation polish or stored recordings on the same connection, Soniox is hard to beat on price. See LecSync vs Soniox.
- Translated speech audio or voice agents. OpenAI's translation and Realtime agent sessions return audio, which LecSync does not. The AssemblyAI and OpenAI comparison goes through what each connection returns.
- Teams already committed to Google Cloud. Savings plans and volume tiers on Speech-to-Text may outweigh list-price differences.
- Live translated captions, speaker labels and running minutes in one product. This is where LecSync is built to win. One WebSocket returns paired transcription and translation, speaker-labelled finals with audio-anchored timestamps, glossary bias and digest sections, and the month is minutes × $0.006. Latency is designed to stay bounded too: server-side backlog is capped at 2 seconds after a network hitch.
Check your own numbers in five steps
- Measure accepted audio minutes and socket-open minutes separately for a week of real traffic.
- List every feature the product actually shows users: captions, translation, speakers, glossary, summary, storage.
- Count target languages per session.
- Price each vendor at its regular rate, then note the promo as upside.
- Convert any second meter (characters, tokens) into dollars per audio hour using your own transcripts, then multiply.
FAQ
What is the cheapest realtime speech-to-text API?
For raw streaming transcription at public list prices on October 5, 2026, Soniox (about $0.12 an hour by its own estimate) and AssemblyAI Universal-Streaming ($0.15 an hour) are among the lowest. Deepgram Nova-3 Monolingual is $0.0048 a minute on promo. LecSync is $0.006 a minute, but that rate already includes translation, speakers and live minutes.
How much do 1,000 hours of realtime transcription cost on LecSync?
1,000 hours is 60,000 minutes, and 60,000 × $0.006 = $360. That figure includes live translation, speaker labels, glossary, live structured summary, AI translation polish and optional 30-day recording storage, according to the public developers page.
Does LecSync charge extra for translation or speaker labels?
No. The developers page lists a single price of $0.006 per audio minute with no tiers or per-feature fees. Disabling features does not reduce the price and enabling all of them does not increase it.
Does LecSync bill for an idle connection?
LecSync bills accepted audio duration, with fractional seconds charged proportionally, and states that no accepted audio means no charge. Each stream also keeps the price snapshot it started with. Session-time billing, as used for AssemblyAI streaming, counts the time a socket is open instead.
What if I need captions in two target languages?
Plan for two meters on most platforms. LecSync configures each stream with one language pair, so a second target language means a second stream at the same rate. OpenAI recommends one translation session per output language. Soniox token costs rise with the extra output text.
Are these prices final?
They are public pay-as-you-go list prices read on October 5, 2026. Enterprise contracts, savings plans and promotions can change them, and Deepgram labels its lower streaming rates as limited-time. Treat the worked examples as a template and rerun them with your own traffic.
The short version
If you need only raw speech-to-text, several vendors undercut $0.006 a minute and you should use them. Once the spec adds live translation, speaker labels and running minutes, add-on stacks grow a second and third meter, and the flat rate starts to look like the simpler and often cheaper bill. Read the LecSync Realtime STT docs for the connect call and the full feature list, then run your own 1,000-hour sheet against it.