Skip to main content
Voice Localization is available through the API only. Listen to each candidate before saving it as a permanent voice.
Choose a flagship voice, a clone or a designed voice, then select a target language and accent. Voice Localization creates candidates for you to preview and save to your voice library.

Starting from a description?

Voice Design creates a new voice from words. Localization edits a voice you already have.

Starting from a recording?

Clone the speaker first, then localize the clone.

How it works

1. Pick a source

A flagship voice, one of your clones, or a Voice Design candidate.

2. Localize

Choose a target language and an accent to generate candidates.

3. Listen

Audition each candidate on a line of Text-to-Speech in the target language.

4. Convert

Save a candidate as a permanent voice in your library.
Each request creates n_samples candidates without modifying the source. A candidate is a temporary voice with a vox_emb_ ID. Preview it with REST Text-to-Speech, then use POST /voices/from-embedding to save it as a permanent voice. The response’s uid is the voice_id you use with REST, WebSocket and Speech-to-Speech. Unsaved candidates expire after 30 days. To change only the accent, keep the source language as target_language and choose another accent. For example, use en and Australian for a British English source.

Access

Base URL https://api.gradium.ai/api, API key in the x-api-key header.

Quickstart

The example localizes a French clone into American English. Set GRADIUM_API_KEY and SOURCE in your environment. SOURCE can be a voice or ready candidate you own, or a flagship voice ID. Shell examples require jq; Python examples require requests. Run the steps in order using one language.
1

List the accents

cURL
200 OK (abbreviated)
Read the list from this endpoint at runtime rather than copying it from these docs: it will grow. You can cache it. The first accent of each language is the default when a request omits accent. Spell accents as listed; matching is case-insensitive and ignores surrounding spaces, but “British”, “Quebecois” or “Québécois French” are rejected.
2

Localize the voice

201 Created
The request is checked before anything is queued, so a bad accent, an unknown source or a source that cannot be localized fails here with 422, 404 or 409 and nothing is billed.
3

Wait for the candidates

200 OK
Localized candidates are typically ready in two to eight seconds. Poll until ready is true before requesting audio. An unknown candidate returns an empty embeddings list. The shell example waits for the first candidate; repeat it for any others you want to preview.kind is edit_language for a localized candidate (generate and enhance for the others). The matching config object, edit_language_config here, contains the request settings, and language is the candidate’s target language. Read the config objects instead of the deprecated prompt field.
4

Listen to a candidate

Pass the candidate id as voice_id on the Text-to-Speech endpoint, with text in the target language.
Compare the candidate with the source reading the same text. Candidate previews use REST only and accept up to 300 characters (400 input text too long beyond). A candidate that is not ready yet returns 404 Embedding not found; an unknown or deleted candidate returns 400.
Use text in the candidate’s target language. Text in another language can produce poor pronunciation. Keep one localized voice per language.
5

Convert the candidate into a voice

Conversion saves the candidate as a permanent voice in your library. Use the response’s uid as voice_id for subsequent speech requests.
201 Created
The saved voice keeps the target language and can be localized again. Conversion is free, uses one custom-voice slot, and removes the candidate’s expiry. Save a candidate you like; a new request produces different results.

Complete example

Set GRADIUM_API_KEY and SOURCE, then run this example to generate candidates and comparison audio in three languages. Listen to the files before choosing which candidates to save. Each saved voice uses a custom-voice slot.
voice_localization.py

Choosing an accent

GET /voice-generator/available-accents returns the accents you can target, grouped by language, with the default first. At the time of writing it covers the five supported languages with between two and six accents each, regional variants such as Quebecois French, Mexican Spanish or Swiss German included. Read the list from the endpoint rather than hardcoding it, and show it to your users as a picker: free-text accents are not accepted. Regional accents the list does not offer, “Bristolian” or “Texan”, belong to Voice Design, where they go in the description.

Gender

Use gender to specify male or female. When you omit it:
  1. A flagship voice uses its own gender tag.
  2. A candidate reuses the gender its own localization was made with.
  3. Otherwise no gender is specified. Clones rarely carry a gender tag, and converted voices do not, so send gender for those sources.
You can request a gender different from the source. Listen to the result to check whether it still matches the speaker you want.

Sources you can localize

Evaluating candidates

  • Compare the source and candidate reading the same target-language text. Listen for pronunciation, accent and whether you still recognize the speaker.
  • For an accent-only comparison, keep the source language and explicitly choose the target accent. The default accent may differ from the source’s accent.
  • Compare multiple candidates before saving one. Repeating a request does not reproduce a previous candidate.
  • Use text in the candidate’s target language. Keep a separate localized voice for each language you need.
  • To clean up a noisy source rather than change its accent, use Voice Enhance.

Candidate lifecycle

Candidates expire after 30 days unless you save them as permanent voices with POST /voices/from-embedding. Saving a candidate removes its expiry and returns the permanent voice’s uid. See the conversion step for the request and response. To remove a candidate early, use DELETE /voice-generator/embeddings/{embedding_id}. A converted voice holds its own copy. Localized candidates also keep a copy of their source, so deleting the source after the request is accepted does not cancel generation. See the Voice Design candidate lifecycle for listing and cleanup details.

Limits

Errors

Errors are {"detail": "..."}, except 422, which carries a list of field errors including the valid accents for the language you asked for.

Next steps

Voice Design

Create the source voice from a description instead.

Custom Voices

Clone a speaker, with a language, so the clone can be localized.

Manage Voices

Set a language on an older clone, list and delete voices.

API Reference

The localization endpoints in full.