Voice Localization is available through the API only.
Listen to each candidate before saving it as a permanent voice.
Starting from a description?
Voice Design creates a new voice from words. Localization edits a voice you already have.
Starting from a recording?
Clone the speaker first, then localize the clone.
How it works
1. Pick a source
A flagship voice, one of your clones, or a Voice Design candidate.
2. Localize
Choose a target language and an accent to generate candidates.
3. Listen
Audition each candidate on a line of Text-to-Speech in the target language.
4. Convert
Save a candidate as a permanent voice in your library.
n_samples candidates without modifying the source.
A candidate is a temporary voice with a vox_emb_ ID. Preview it with REST
Text-to-Speech, then use POST /voices/from-embedding to save it as a permanent
voice. The response’s uid is the voice_id you use with REST, WebSocket and
Speech-to-Speech. Unsaved candidates expire after 30 days.
To change only the accent, keep the source language as target_language and
choose another accent. For example, use en and Australian for a British
English source.
Access
Base URLhttps://api.gradium.ai/api, API key in the x-api-key header.
Quickstart
The example localizes a French clone into American English. SetGRADIUM_API_KEY and SOURCE in your environment. SOURCE can be a voice or
ready candidate you own, or a flagship voice ID. Shell examples require jq;
Python examples require requests. Run the steps in order using one language.
1
List the accents
cURL
200 OK (abbreviated)
accent. Spell accents as
listed; matching is case-insensitive and ignores surrounding spaces, but
“British”, “Quebecois” or “Québécois French” are rejected.2
Localize the voice
201 Created
The request is checked before anything is queued, so a bad accent, an
unknown source or a source that cannot be localized fails here with
422,
404 or 409 and nothing is billed.3
Wait for the candidates
200 OK
ready is true before requesting audio. An unknown candidate
returns an empty embeddings list. The shell example waits for the first
candidate; repeat it for any others you want to preview.kind is edit_language for a localized candidate (generate and enhance
for the others). The matching config object, edit_language_config here,
contains the request settings, and language is the candidate’s target
language. Read the config objects instead of the deprecated prompt field.4
Listen to a candidate
Pass the candidate id as Compare the candidate with the source reading the same text. Candidate
previews use REST only and accept up to 300 characters (
voice_id on the Text-to-Speech endpoint, with
text in the target language.400 input text too long beyond). A candidate that is not ready yet returns
404 Embedding not found; an unknown or deleted candidate returns 400.5
Convert the candidate into a voice
Conversion saves the candidate as a permanent voice in your library.
Use the response’s The saved voice keeps the target
uid as voice_id for subsequent speech requests.201 Created
language and can be localized again.
Conversion is free, uses one custom-voice slot, and removes the candidate’s
expiry. Save a candidate you like; a new request produces different results.Complete example
SetGRADIUM_API_KEY and SOURCE, then run this example to generate candidates
and comparison audio in three languages. Listen to the files before choosing
which candidates to save. Each saved voice uses a custom-voice slot.
voice_localization.py
Choosing an accent
GET /voice-generator/available-accents returns the accents you can target,
grouped by language, with the default first. At the time of writing it covers
the five supported languages with between two and six accents each, regional
variants such as Quebecois French, Mexican Spanish or Swiss German included.
Read the list from the endpoint rather than hardcoding it, and show it to your
users as a picker: free-text accents are not accepted. Regional accents the
list does not offer, “Bristolian” or “Texan”, belong to
Voice Design, where they go in the description.
Gender
Usegender to specify male or female. When you omit it:
- A flagship voice uses its own gender tag.
- A candidate reuses the gender its own localization was made with.
- Otherwise no gender is specified. Clones rarely carry a gender tag,
and converted voices do not, so send
genderfor those sources.
Sources you can localize
Evaluating candidates
- Compare the source and candidate reading the same target-language text. Listen for pronunciation, accent and whether you still recognize the speaker.
- For an accent-only comparison, keep the source language and explicitly choose the target accent. The default accent may differ from the source’s accent.
- Compare multiple candidates before saving one. Repeating a request does not reproduce a previous candidate.
- Use text in the candidate’s target language. Keep a separate localized voice for each language you need.
- To clean up a noisy source rather than change its accent, use Voice Enhance.
Candidate lifecycle
Candidates expire after 30 days unless you save them as permanent voices withPOST /voices/from-embedding. Saving a candidate removes its expiry and returns
the permanent voice’s uid. See the conversion step for the request
and response.
To remove a candidate early, use
DELETE /voice-generator/embeddings/{embedding_id}. A converted voice holds its
own copy. Localized candidates also keep a copy of their source, so deleting
the source after the request is accepted does not cancel generation. See the
Voice Design candidate lifecycle
for listing and cleanup details.
Limits
Errors
Errors are{"detail": "..."}, except 422, which carries a list of field
errors including the valid accents for the language you asked for.
Next steps
Voice Design
Create the source voice from a description instead.
Custom Voices
Clone a speaker, with a language, so the clone can be localized.
Manage Voices
Set a language on an older clone, list and delete voices.
API Reference
The localization endpoints in full.