Realtime Speech API

LecSync
Realtime STT

  • General-purpose transcription model
  • Fast transcription
  • Realtime translation
  • Speaker diarization

A single WebSocket connection providing realtime transcription and translation. The first word reaches the screen 0.28 seconds after someone starts speaking.

Sample session replaylanguage auto-detected
Speaker 1ja

零点付近の量子化ステップは、およそ 8 です。

The quantization step near zero is about eight.

Speaker 2de

Dann bleiben ungefähr 22 dB Headroom, oder?

That leaves roughly 22 dB of headroom, right?

Speaker 1ko

맞아요 — 노이즈가 바로 거기서 올라오기 시작합니다.

Right — that is exactly where the noise starts coming up.

into

Select a timestamp to jump to that line. This is a replay of a sample session, not the live API.

60
languages, translated between any two
0.28 s
from speech to first word on screen
5
global gateway regions
$0.01
per audio minute, all features included
01Realtime translation

Direction is determined by the content

Each stream is configured with a single language pair, and the direction is determined per segment from the recognition result.

If a third language appears during the session, it is translated into your target language as well. Translation streams alongside the transcript rather than waiting for the sentence to complete.

Translation adds approximately 88 ms to the first token.

60

languages, translated between any two

English中文日本語한국어EspañolFrançaisDeutschItalianoPortuguêsРусскийहिन्दीالعربيةNederlandsPolskiTürkçeУкраїнськаČeštinaSlovenčinaSlovenščinaMagyarRomânăБългарскиHrvatskiСрпскиМакедонскиΕλληνικάSvenskaDanskNorskSuomiEestiLatviešuLietuviųShqipCatalàAfrikaansעבריתفارسیاردوमराठीবাংলাગુજરાતીಕನ್ನಡമലയാളംਪੰਜਾਬੀதமிழ்తెలుగుไทยTiếng ViệtBahasa IndonesiaBahasa MelayuTagalogKiswahiliҚазақшаAzərbaycanEuskaraБеларускаяBosanskiGalegoCymraeg
02Realtime transcription

Text appears before the sentence ends

The first word reaches the screen 0.28 seconds after a speaker begins.

Preview text streams continuously during recognition and settles into a final segment once confirmed. Language is detected automatically, with no need to select it before the session starts.

03Latency control

Latency does not accumulate

When the network slows, audio builds up server-side. We cap that backlog at 2 seconds and discard anything beyond it rather than queueing it.

Once the disruption passes, captions return to the current sentence within a second instead of falling progressively further behind.

04Speaker diarization

Every segment is labelled with a speaker

With speakers enabled, every final segment carries a speaker label.

Labels are returned together with timestamps, so a given speaker's turns can be located directly in the audio.

05Audio alignment

Timestamps anchored to the audio

A segment's startMs is its position within the recording and can be used for playback directly, with no further correction.

The upstream recognition clock counts within a single connection and resets on reconnect. The gateway converts it into a global recording time using the byte offset of the audio file.

06Live minutes

Minutes build up section by section

Pass digest: true and a section of minutes is produced once enough has been said, pushed over the same connection.

When a topic finishes, its section is marked confirmed; the current topic stays open and keeps gathering points. Each section carries a time range that maps back onto the recording.

07Global infrastructure

The same API, deployed across 5 regions

Audio is processed close to where it is captured, so frames are not carried across an ocean. Specify a region during the handshake, or let us route to the nearest healthy node.

DAL

region: "dal"

HZ

region: "hz"

JP2

region: "jp2"

Russia

region: "ru"

United Kingdom

region: "uk"

08Additional capabilities

All within the same connection

No additional endpoint, no additional integration, and no additional line on the invoice.

Glossary & domain context

Biases recognition toward your terminology and fixes the translation of names and product names.

AI translation polish

A second model pass rewrites the translation with surrounding context and republishes it under the same segment ID.

Recording storage

store defaults to false. Pass store: true to keep the recording and transcript JSON for 30 days after the stream ends, retrieved with a short-lived signed URL.

09API scope

Handled on our side

Audio handling, time alignment, terminology, minutes and flow control are all resolved server-side.

Audio buffering, encoding, upload, storage

Clients push raw PCM after session.ready. store defaults to false; nothing is kept unless you opt in.

Aligning transcript to audio

Segment timestamps correspond directly to a position in the recording.

Stitching transcription to translation

Transcription and translation are returned over the same connection, already paired.

Glossary injection

Product names and specialist terms remain consistent across source and translation.

The summarization pipeline

Structured minutes are pushed continuously while the session runs.

Flow control and backpressure

Congestion is resolved server-side; the client simply keeps sending.

GPU operations

Translation runs on models we fine-tuned, on hardware we operate.

10Integration

Connect, then stream

Connect with curl, then use its complete url and token in either short example to send 3 seconds of silence. The docs include complete WAV clients with bounded queues, timeouts and cleanup.

curl — create a session

curl -X POST "https://api.lecsync.com/v1/realtime/connect" \
  -H "Authorization: Bearer $LECSYNC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"audio":{"encoding":"pcm_s16le","sampleRate":16000,"channels":1},
       "transcribe":{"languages":["en"]},"store":false}'
11Pricing

A single price

No tiers or per-feature fees. Each stream is billed in whole minutes, rounded up.

$0.01per minute of audio

All of the following are included:

  • Realtime transcription
  • Real-time translation
  • Audio-aligned timestamps
  • Speaker separation
  • Glossary and domain context
  • AI translation refinement
  • Live structured summary
  • Recording + transcript storage, 30 days

Each stream is rounded up to whole audio minutes at settlement. Reconnecting starts a separately billed stream. No accepted audio means no charge.

Disabling features does not reduce the price, and enabling all of them does not increase it.

If a translation node becomes unreachable, the request falls back automatically and the stream continues.

Open your first stream in minutes

Register, create an API key, and start with the two calls above.