Browse documentation

Timestamps

startMs and endMs on transcript.final and translation.final are millisecond offsets into this stream's audio, measured from the first byte you sent. With store: true they are positions in the recording you can download.

Recognisers report timing relative to their own upstream connection, and that connection re-establishes mid-session without telling you. When it does, the clock resets to zero and click-to-play starts jumping backwards. Short tests never reconnect, so this surfaces in production.

Timestamps come from a byte clock over the audio you sent, not from the recogniser's own clock, so they stay monotonic across upstream reconnects. A small number of segments may arrive without timestamps when the recogniser does not supply token-level timing.

With store: true, one connection produces one recording. If your network drops mid-session you get multiple files — order them by creation time.