Read the spoken segment
The input is an ASR chunk: colloquial, cut off, or slightly wrong. The model translates what is there instead of repairing it into an essay.
Our 4B translation model reads live lecture and meeting transcripts — with prior context and your glossary — and writes only the translation. It is the engine behind LecSync live translation. The lighter 1.7B is still available.
Developer API coming soon
In the product
LecSync-MT is not a side research demo. It is the translation model LecSync runs on live speech: one ASR segment in, optional previous segments and a course glossary beside it, translation only out. That contract matches how lectures and meetings actually unfold — short, spoken, sometimes incomplete.
How it works
Generic MT models are trained on written documents. LecSync-MT is trained on speech transcripts, so it keeps fragments as fragments and does not invent the rest of the sentence.
The input is an ASR chunk: colloquial, cut off, or slightly wrong. The model translates what is there instead of repairing it into an essay.
Previous source-language segments are optional context. They resolve pronouns and topic words without being translated a second time.
Pass course or product terms as source → target pairs. The model is supervised to keep those forms instead of a fluent near-miss.
[system] You are a professional simultaneous interpreter for live lectures and meetings. [user] Context (previous segments): Today we talk about database design. Glossary: ERD -> 实体关系图 English->Chinese: So we use the ERD first.
Evaluations
Quality calls use a pairwise LLM judge (Gemini, both orders). Official BLEU is published as a reference number only — it disagreed with judged quality on three frozen sets.
Judge: Gemini 3.8 Flash, both orders, on frozen lecture and meeting sets. The production comparison uses 407 segments from live LecSync traffic. Unless noted, results below are for the 1.7B.
4B vs 1.7B
4.21 / 5
Half the serious errors
Blind 1–5 scoring on 998 live segments: 4B averages 4.21 vs 3.72 for 1.7B, with serious errors down from 18.2% to 9.9%. Judged best on 485 segments vs 58.
Previous cloud pipeline
51.9%
Statistically tied
Win rate against the previous production stack (cloud draft plus optional rewrite) on 407 lecture and meeting segments. p = 0.69. Against the raw cloud draft alone the point estimate is 55.7%, still not significant.
Tencent Hy-MT 1.8B
71.9%
Ahead on lecture ASR
Win rate on general lecture segments. With prior context: 74.3%. With a glossary: 69.3%. Hy-MT leads on WMT24 written and news text — LecSync-MT is a speech specialist, not a general literary model.
Qwen3.5-2B
6 / 6
Ahead on every frozen set
Same six frozen suites as the Hy-MT comparison, including lecture, context, glossary, no-context, WMT24, and multilingual targets.
| Suite | LecSync-MT win rate |
|---|---|
| Lecture segments | 71.9% |
| Needs prior context | 74.3% |
| Glossary terms | 69.3% |
Specs
Task
Speech-transcript translation
Training
Real lecture and meeting transcripts
Training rows
149,042
Decoding
Greedy (as evaluated)
Most lecture rows come from real LecSync transcripts, labelled by a teacher model and filtered by an LLM judge. A multilingual replay mix keeps the original 60 target-language slot alive so the model does not collapse every target to Chinese. Context-sensitive pairs and glossary rows are in the mix; replay rows carry neither.
Ask for a target with English language names (English→Chinese, Korean→English). Coverage is strongest on the lecture languages we serve in the product. A few low-resource scripts have thinner quality.
The hosted playground lets you send a lecture or meeting segment and see LecSync-MT translate it. Weights are not released. A public HTTP API is not available yet.
The fastest way to judge LecSync-MT is to open LecSync and translate a real lecture or meeting.