Speakers
Turn diarization on at handshake and every settled segment comes back with a speaker label next to its timestamps.
Settled segments carry a speaker field; partial results do not, because attribution is only decided once a segment settles.
The speaker label arrives on the same message as startMs and endMs, which makes a transcript playable per speaker: pick a label, seek to the offset, press play.