json_config
parameter. In the Python SDK, this is a dict mapping option name to a
JSON-serializable value. When using the REST endpoint, pass it as a
URL-encoded JSON string in the json_config query parameter.
These options apply to both the WebSocket
and REST transports. For TTS settings,
see Voice Settings.
Quick reference
Default values may evolve with new model releases. Pin the options
explicitly if you depend on a specific behaviour. The server validates
known keys with the constraints above; unknown keys are silently
ignored, so double-check spelling against the table.
Languages
For non-translating transcription models,language is the expected language
of the audio. It grounds the model to a single language for better transcription
quality. For audio that mixes languages, set the dominant one — or leave the
option unset when no single language predominates. target_language has no
effect on these models.
For translating transcription models, language and target_language specify
the language the transcription should be generated in, not the language of the
audio. The two are interchangeable; set either one.
Keyword boosting
Adapt transcription to your own vocabulary: player names, brands, products, and channel names. Pass akeywords dictionary and
the decoder gives those terms priority, so the vocabulary that matters
to your product is transcribed exactly as you expect, in real time and
without retraining.
words: the terms to boost. Matching is token-by-token, so keywords must not contain spaces. Split multi-word names, so"Ferran Torres"becomes"Ferran"and"Torres". Matching is also case and accent sensitive, so include the variants you expect ("Mbappé"and"Mbappe"). You can boost up to 500 keywords, and every variant of a word (lowercase, uppercase, accented, etc.) counts toward that limit.boost: how strongly to bias toward the dictionary, from-6to6. The value is applied in log-probability space, so its effect is exponential: small values already shift decoding noticeably.
3 is the recommended default and works for most
cases; raise it to 4 for unusually rare or foreign vocabulary. Higher
values give diminishing returns and can make the decoder loop,
repeating a boosted word instead of following the audio, so avoid going
much beyond 4 to 5.
Negative values de-boost a term. Use them only to remove a
specific word you never want in the output, not as a general accuracy
control.
Passing json_config
The same json_config payload is sent regardless of which SDK API you
use; only the call shape differs:
json_config as a
URL-encoded JSON string in the query parameters, see the
REST guide.
Semantic VAD and delay
delay_in_frames is most useful with the WebSocket STT stream. The
server emits semantic VAD step messages every 80 ms; each step
contains future horizons with inactivity_prob values. A common
starting point for voice agents is to watch the longest horizon and
flush when the inactivity probability stays above your threshold.
delay_in_frames values for fast back-and-forth assistants,
and higher values for transcription quality when a little more latency
is acceptable. A value such as 16 is a balanced starting point; try
8 when responsiveness matters more than context.