Back to Blog

Built-in Translation vs DIY STT + MT for Realtime Speech APIs

LecSync Team

Teams that already run machine translation often ask a practical question: should a realtime speech API with built-in translation handle live captions end to end, or should you keep STT and MT as two services you stitch yourself?

Quick verdict: A DIY STT + MT pipeline is a fair choice when you only need raw transcripts and already own a translation stack you trust. Choose a single WebSocket that pairs transcription and translation server-side—such as LecSync Realtime STT—when live bilingual captions are the product: one bill at $0.006 per audio minute, speakers, digest, and glossary on the same connection, a 2-second backlog cap, and five gateway regions.

What is a realtime speech API with built-in translation?

A realtime speech API with built-in translation accepts live audio over one streaming connection and returns source transcripts paired with target-language text as speech continues—without a second vendor call or a client-side join step. On LecSync’s developers page, that means one WebSocket for transcription and translation, with translation adding about 88 ms to the first translation token after the first caption word appears in roughly 0.28 seconds.

When DIY STT + MT still makes sense

Stitching a streaming STT socket to a separate MT API is not a mistake by default. It fits when:

  • Your product only shows source-language captions and translation is optional or offline
  • You already pay for (and tune) an in-house or contracted MT engine for documents, chat, or CMS copy
  • Latency budgets are loose enough that a second hop after each final segment is acceptable
  • You want to swap STT or MT vendors independently for experiments

In those cases, keep the STT vendor focused on recognition quality and let your existing MT path own terminology and style. The rest of this article is about the failure modes that show up once you need live source + target lines that stay aligned under network stress.

Side-by-side: built-in pairing vs a DIY pipeline

Figures for LecSync were rechecked on the public developers page ~2026-10-02 (Asia/Shanghai). DIY rows describe a common architecture pattern—separate realtime STT WebSocket plus a separate MT HTTP/stream API—not a named competitor’s price sheet.

DimensionBuilt-in (LecSync one WebSocket)DIY STT + MT (typical stitch)
Connection modelOne WebSocket; transcription + translation paired server-sideSTT stream + separate MT calls; client or middleware joins IDs
Languages60, any-to-any pairs on the same sessionDepends on each vendor’s language lists and overlap
First word / translation lag0.28 s to first word; translation ~+88 ms to first translation tokenSTT latency + MT queue + your join logic
Speakers / timestampsLabels and audio-aligned timestamps on finalsSTT may provide them; MT usually does not—you keep mapping
Digest / glossarydigest:true and glossary on the same connectionExtra services or custom jobs after the fact
Backlog under hitchCap 2 s; captions return to the current sentence within ~1 sTwo queues can grow independently; catch-up is your problem
Regions5 gateways: dal, hz, jp2, ru, ukEach vendor’s regions; audio may cross oceans twice
Pricing shapeFlat $0.006 / audio minute, listed features includedSTT meter + MT meter (+ retries, polish, storage if separate)

Sources: LecSync developers. DIY column is architectural, not a third-party list price.

Failure modes of stitching two vendors

Two meters, two invoices. Finance asks what 10,000 live caption-minutes cost. With LecSync the answer is multiplication: minutes × $0.006. With DIY you add STT duration (or tokens) to MT characters/tokens, then decide who pays for partial segments, empty retries, and a second model pass for polish.

Desync under partial failure. STT finals arrive while MT times out; the UI shows orphan source lines or stale translations until you invent backoff and segment ID rules. A built-in path returns paired text on one socket, so the product surface stays one stream of events.

Double retries, double backpressure. A network hitch can stall audio into the STT vendor and MT requests into a second queue. LecSync documents a 2-second server-side backlog cap: audio beyond the cap is discarded rather than queued forever, and captions catch the current sentence within about a second after recovery. In DIY, both vendors—and your glue—need an equivalent policy or latency grows without a ceiling.

Feature sprawl outside the caption path. Speaker labels, glossary bias, live structured minutes (digest), and optional 30-day recording storage sit on LecSync’s same connection at the same rate. DIY often means a third summarizer, a custom glossary store, and storage you operate yourself.

Pricing: one all-in minute vs two line items

LecSync publishes a single public rate: $0.006 per minute of audio. Included on that rate: realtime transcription, realtime translation, audio-aligned timestamps, speaker separation, glossary and domain context, AI translation polish, live structured summary, and optional 30-day recording/transcript storage. Disabling features does not lower the price; enabling all of them does not raise it. Fractional seconds are charged proportionally; no accepted audio means no charge.

A DIY spreadsheet needs at least two vendor rows plus engineering time for join logic. Exact STT and MT list prices vary by vendor and plan—verify each public pricing page for the services you shortlist. The selection question is whether you want one forecastable audio-minute line or a sum that moves whenever either side changes metering.

Orientation example (LecSync public list only)

StackHow cost is framed publicly60 minutes of accepted audio
LecSync all-in$0.006 / audio minute$0.36
DIY STT + MTSTT duration/tokens + MT volume (+ glue)Varies by vendors—sum both invoices

When to pick DIY STT + MT

Pick the stitch when:

  • Live translation is rare or batch-only
  • Your MT quality/glossary already beats what a speech vendor bundles
  • You have staffed SRE time for two SLAs, two dashboards, and join-edge cases
  • You are prototyping recognition quality and will not ship bilingual UI yet

When to pick a built-in realtime speech API (LecSync)

Pick LecSync when the brief says ship live source + target captions as a product:

  • Classroom, webinar, or meeting clients that show both lines as people speak
  • Apps that need speaker-labelled finals and structured digest minutes without a second summarization vendor
  • Teams that want one invoice line at $0.006 / audio minute
  • Launches where desync between STT and MT would become a support queue

Connect with POST https://api.lecsync.com/v1/realtime/connect and a Bearer key, then stream on the returned WebSocket (developers). Preview text streams, then finals; language auto-detect is available. Translation streams alongside the transcript rather than waiting for the full sentence.

Integration scenarios

Live bilingual captions. DIY means every final segment triggers an MT call and a UI merge. Built-in means the client renders paired events from one socket.

Forecastable COGS. Minutes × $0.006 is a finance-ready formula. Two vendors need a model for MT volume growth when speakers talk faster or when you enable polish.

Recovery after a hitch. One backlog policy (2 s cap) is easier to reason about than STT buffer + MT retry storm + your own queue.

FAQ

Is DIY STT + MT always cheaper than a built-in speech API?

Not automatically. DIY can win if you already pay for MT at scale and only need occasional translation of finals. Built-in wins on predictability when every live minute needs a paired translation and you would otherwise pay two meters plus glue. LecSync’s public all-in rate is $0.006 per audio minute.

What does “paired server-side” mean?

Transcription and translation are produced and returned on the same WebSocket, already associated with the same segment identity. The client does not call a second API to translate each final line.

How fast is LecSync’s first caption word?

LecSync publishes 0.28 seconds from speech start to first word on screen, and about 88 ms added to the first translation token. Treat those as vendor product figures.

Does LecSync charge extra for translation, speakers, or digest?

No. The public rate stays $0.006 per audio minute. Listed features are included; turning them off does not discount the minute.

How many languages and regions does LecSync cover?

60 languages with any-to-any translation pairs, and 5 gateway regions (dal, hz, jp2, ru, uk) on the public developers page.

What happens to latency when the network stalls?

LecSync caps server-side audio backlog at 2 seconds and discards audio beyond the cap. After recovery, captions return to the current sentence within about a second instead of falling further behind.

Verdict

DIY STT + MT remains reasonable when translation is secondary and you already operate a strong MT path. A built-in realtime speech API fits when live bilingual (or any-to-any) captions are the product and you want transcription, translation, speakers, glossary, and digest on one WebSocket at a flat $0.006 per audio minute, with a documented 2-second backlog cap and five regions. Start at the LecSync Realtime STT docs and compare that single line item against the sum of your STT vendor, MT vendor, and the glue you would maintain.