> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gradium.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice Localization

> Adapt an existing voice to another language or accent

<Note>
  Voice Localization is available through the API only.
  Listen to each candidate before saving it as a permanent voice.
</Note>

Choose a flagship voice, a clone or a designed voice, then select a target
language and accent. Voice Localization creates candidates for you to preview
and save to your voice library.

<CardGroup cols={2}>
  <Card title="Starting from a description?" icon="wand-magic-sparkles" href="/guides/voices/voice-design">
    Voice Design creates a new voice from words. Localization edits a voice you already have.
  </Card>

  <Card title="Starting from a recording?" icon="microphone" href="/guides/voices/custom-voices">
    Clone the speaker first, then localize the clone.
  </Card>
</CardGroup>

## How it works

<CardGroup cols={2}>
  <Card title="1. Pick a source" icon="user">
    A flagship voice, one of your clones, or a Voice Design candidate.
  </Card>

  <Card title="2. Localize" icon="language">
    Choose a target language and an accent to generate candidates.
  </Card>

  <Card title="3. Listen" icon="headphones">
    Audition each candidate on a line of Text-to-Speech in the target language.
  </Card>

  <Card title="4. Convert" icon="bookmark">
    Save a candidate as a permanent voice in your library.
  </Card>
</CardGroup>

Each request creates `n_samples` candidates without modifying the source.
A candidate is a temporary voice with a `vox_emb_` ID. Preview it with REST
Text-to-Speech, then use `POST /voices/from-embedding` to save it as a permanent
voice. The response's `uid` is the `voice_id` you use with REST, WebSocket and
Speech-to-Speech. Unsaved candidates expire after 30 days.

To change only the accent, keep the source language as `target_language` and
choose another accent. For example, use `en` and `Australian` for a British
English source.

## Access

Base URL `https://api.gradium.ai/api`, API key in the `x-api-key` header.

| Call | Endpoint |
| :- | :- |
| List the accents you can target | `GET /voice-generator/available-accents` |
| Localize a voice | `POST /voice-generator/localize` |
| Check readiness, or list candidates | `GET /voice-generator/embeddings` |
| Audition a candidate | `POST /post/speech/tts` |
| Convert a candidate | `POST /voices/from-embedding` |
| Delete a candidate | `DELETE /voice-generator/embeddings/{embedding_id}` |

## Quickstart

The example localizes a French clone into American English. Set
`GRADIUM_API_KEY` and `SOURCE` in your environment. `SOURCE` can be a voice or
ready candidate you own, or a flagship voice ID. Shell examples require `jq`;
Python examples require `requests`. Run the steps in order using one language.

<Steps>
  <Step title="List the accents">
    ```bash cURL theme={null}
    curl --fail --silent --show-error --max-time 30 -H "x-api-key: $GRADIUM_API_KEY" \
      https://api.gradium.ai/api/voice-generator/available-accents
    ```

    ```json 200 OK (abbreviated) theme={null}
    {
      "languages": {
        "en": ["General American", "Standard British", "Irish", "..."],
        "fr": ["General French", "Quebecois French"],
        "...": ["..."]
      }
    }
    ```

    Read the list from this endpoint at runtime rather than copying it from
    these docs: it will grow. You can cache it. The first accent of each
    language is the default when a request omits `accent`. Spell accents as
    listed; matching is case-insensitive and ignores surrounding spaces, but
    "British", "Quebecois" or "Québécois French" are rejected.
  </Step>

  <Step title="Localize the voice">
    <CodeGroup>
      ```bash cURL theme={null}
      curl --fail --silent --show-error --max-time 30 -X POST https://api.gradium.ai/api/voice-generator/localize \
        -H "x-api-key: $GRADIUM_API_KEY" \
        -H "Content-Type: application/json" \
        -d "{
          \"src_voice\": \"$SOURCE\",
          \"target_language\": \"en\",
          \"accent\": \"General American\",
          \"gender\": \"female\",
          \"n_samples\": 2
        }" > candidates.json

      CAND0=$(jq -r '.embeddings[0].embedding_id' candidates.json)
      ```

      ```python Python theme={null}
      import os
      import requests

      API_KEY = os.environ["GRADIUM_API_KEY"]
      BASE = "https://api.gradium.ai/api"
      SOURCE = os.environ["SOURCE"]  # a voice id or a vox_emb_ candidate id

      resp = requests.post(
          f"{BASE}/voice-generator/localize",
          json={
              "src_voice": SOURCE,
              "target_language": "en",
              "accent": "General American",
              "gender": "female",
              "n_samples": 2,
          },
          headers={"x-api-key": API_KEY},
          timeout=30,
      )
      resp.raise_for_status()
      candidates = [e["embedding_id"] for e in resp.json()["embeddings"]]
      ```
    </CodeGroup>

    ```json 201 Created theme={null}
    {
      "embeddings": [
        { "embedding_id": "vox_emb_XeBcXDgoigTfL9Cv", "ready": false, "expires_at": "2026-10-21T13:03:52Z" },
        { "embedding_id": "vox_emb_ni8lehVHTb35siXI", "ready": false, "expires_at": "2026-10-21T13:03:53Z" }
      ]
    }
    ```

    | Field | Required | Description |
    | :- | :- | :- |
    | `src_voice` | required | A voice id (your clone, a converted candidate, or a flagship voice) or a `vox_emb_` candidate id. |
    | `target_language` | required | `en`, `fr`, `es`, `pt` or `de`. May equal the source language for an accent change. |
    | `accent` | optional | One of the accents listed for `target_language`. Defaults to the first one. |
    | `gender` | optional | `male` or `female`. See [Gender](#gender). |
    | `n_samples` | optional | 1 to 5 candidates, default 1. |

    The request is checked before anything is queued, so a bad accent, an
    unknown source or a source that cannot be localized fails here with `422`,
    `404` or `409` and nothing is billed.
  </Step>

  <Step title="Wait for the candidates">
    <CodeGroup>
      ```bash cURL theme={null}
      wait_for_candidate() {
        local deadline=$((SECONDS + 120))
        local result
        while (( SECONDS < deadline )); do
          result=$(curl --fail --silent --show-error --max-time 10 \
            -H "x-api-key: $GRADIUM_API_KEY" \
            "https://api.gradium.ai/api/voice-generator/embeddings?embedding_id=$1") || return 1
          if ! jq -e '.embeddings | length > 0' <<< "$result" > /dev/null; then
            echo "Candidate not found: $1" >&2
            return 1
          fi
          if jq -e '.embeddings[0].ready' <<< "$result" > /dev/null; then
            return 0
          fi
          sleep 2
        done
        echo "Timed out waiting for candidate: $1" >&2
        return 1
      }
      wait_for_candidate "$CAND0"
      ```

      ```python Python theme={null}
      import time

      def wait_until_ready(embedding_id, timeout_s=120):
          deadline = time.monotonic() + timeout_s
          while time.monotonic() < deadline:
              resp = requests.get(
                  f"{BASE}/voice-generator/embeddings",
                  params={"embedding_id": embedding_id},
                  headers={"x-api-key": API_KEY},
                  timeout=min(10, max(0.1, deadline - time.monotonic())),
              )
              resp.raise_for_status()
              found = resp.json()["embeddings"]
              if not found:
                  raise RuntimeError(f"candidate {embedding_id} not found")
              if found and found[0]["ready"]:
                  return found[0]
              time.sleep(2)
          raise TimeoutError(embedding_id)

      for embedding_id in candidates:
          wait_until_ready(embedding_id)
      ```
    </CodeGroup>

    ```json 200 OK theme={null}
    {
      "embeddings": [{
        "embedding_id": "vox_emb_XeBcXDgoigTfL9Cv",
        "ready": true,
        "kind": "edit_language",
        "language": "en",
        "created_at": "2026-09-21T13:03:52.591128",
        "expires_at": "2026-10-21T13:03:52.495534",
        "generate_config": null,
        "edit_language_config": {
          "src_voice": "H5oin0KHqRTAxaeL",
          "src_language": "fr",
          "target_language": "en",
          "accent": "General American",
          "gender": "female"
        },
        "enhance_config": null,
        "prompt": "Feminine. General American accent."
      }]
    }
    ```

    Localized candidates are typically ready in two to eight seconds. Poll
    until `ready` is `true` before requesting audio. An unknown candidate
    returns an empty `embeddings` list. The shell example waits for the first
    candidate; repeat it for any others you want to preview.

    `kind` is `edit_language` for a localized candidate (`generate` and `enhance`
    for the others). The matching config object, `edit_language_config` here,
    contains the request settings, and `language` is the candidate's target
    language. Read the config objects instead of the deprecated `prompt` field.
  </Step>

  <Step title="Listen to a candidate">
    Pass the candidate id as `voice_id` on the Text-to-Speech endpoint, with
    text **in the target language**.

    <CodeGroup>
      ```bash cURL theme={null}
      curl --fail --silent --show-error --max-time 30 -X POST https://api.gradium.ai/api/post/speech/tts \
        -H "x-api-key: $GRADIUM_API_KEY" \
        -H "Content-Type: application/json" \
        -d "{
          \"text\": \"Hi, this is Constance. Same voice, new language, and hopefully the same personality.\",
          \"voice_id\": \"$CAND0\",
          \"model_name\": \"default\",
          \"output_format\": \"wav\",
          \"only_audio\": true
        }" --output candidate-0.wav
      ```

      ```python Python theme={null}
      resp = requests.post(
          f"{BASE}/post/speech/tts",
          json={
              "text": "Hi, this is Constance. Same voice, new language, and hopefully the same personality.",
              "voice_id": candidates[0],
              "model_name": "default",
              "output_format": "wav",
              "only_audio": True,
          },
          headers={"x-api-key": API_KEY},
          timeout=30,
      )
      resp.raise_for_status()
      with open("candidate-0.wav", "wb") as f:
          f.write(resp.content)
      ```
    </CodeGroup>

    Compare the candidate with the source reading the same text. Candidate
    previews use REST only and accept up to 300 characters (`400 input text
            too long` beyond). A candidate that is not ready yet returns
    `404 Embedding not found`; an unknown or deleted candidate returns `400`.

    <Warning>
      Use text in the candidate's target language. Text in another language can
      produce poor pronunciation. Keep one localized voice per language.
    </Warning>
  </Step>

  <Step title="Convert the candidate into a voice">
    Conversion saves the candidate as a permanent voice in your library.
    Use the response's `uid` as `voice_id` for subsequent speech requests.

    <CodeGroup>
      ```bash cURL theme={null}
      curl --fail --silent --show-error --max-time 30 -X POST https://api.gradium.ai/api/voices/from-embedding \
        -H "x-api-key: $GRADIUM_API_KEY" \
        -H "Content-Type: application/json" \
        -d "{
          \"voxium_embedding_id\": \"$CAND0\",
          \"name\": \"Constance EN\",
          \"description\": \"French support voice, localized to General American\"
        }" > voice.json

      VOICE=$(jq -r '.uid' voice.json)
      ```

      ```python Python theme={null}
      resp = requests.post(
          f"{BASE}/voices/from-embedding",
          json={
              "voxium_embedding_id": candidates[0],
              "name": "Constance EN",
              "description": "French support voice, localized to General American",
          },
          headers={"x-api-key": API_KEY},
          timeout=30,
      )
      resp.raise_for_status()
      voice_id = resp.json()["uid"]
      ```
    </CodeGroup>

    ```json 201 Created theme={null}
    {
      "uid": "w0F73K2fcf3X2Vcn",
      "name": "Constance EN",
      "description": "French support voice, localized to General American",
      "filename": "vox_emb_XeBcXDgoigTfL9Cv",
      "start_s": 0.0,
      "is_catalog": false,
      "is_pro_clone": false,
      "language": "en",
      "tags": []
    }
    ```

    The saved voice keeps the target `language` and can be localized again.
    Conversion is free, uses one custom-voice slot, and removes the candidate's
    expiry. Save a candidate you like; a new request produces different results.
  </Step>
</Steps>

## Complete example

Set `GRADIUM_API_KEY` and `SOURCE`, then run this example to generate candidates
and comparison audio in three languages. Listen to the files before choosing
which candidates to save. Each saved voice uses a custom-voice slot.

```python voice_localization.py theme={null}
import os
import time

import requests

BASE_URL = "https://api.gradium.ai/api"
HEADERS = {"x-api-key": os.environ["GRADIUM_API_KEY"]}
SOURCE = os.environ["SOURCE"]

TARGETS = [  # (language, accent, audition line)
    ("en", "General American", "Hi, this is your assistant. How can I help you today?"),
    ("de", "Standard German", "Hallo, hier ist Ihre Assistentin. Wie kann ich Ihnen heute helfen?"),
    ("es", "Mexican Spanish", "Hola, soy tu asistente. ¿En qué puedo ayudarte hoy?"),
]


def check(resp):
    if not resp.ok:
        raise RuntimeError(f"HTTP {resp.status_code}: {resp.text}")
    return resp


def localize(src_voice, language, accent, gender=None, n_samples=1):
    body = {"src_voice": src_voice, "target_language": language, "accent": accent, "n_samples": n_samples}
    if gender:
        body["gender"] = gender
    resp = check(requests.post(f"{BASE_URL}/voice-generator/localize", headers=HEADERS, json=body, timeout=30))
    return [c["embedding_id"] for c in resp.json()["embeddings"]]


def wait_until_ready(candidate_ids, timeout_s=120.0):
    deadline = time.monotonic() + timeout_s
    pending = set(candidate_ids)
    while pending:
        for candidate_id in sorted(pending):
            resp = check(requests.get(
                f"{BASE_URL}/voice-generator/embeddings",
                headers=HEADERS,
                params={"embedding_id": candidate_id},
                timeout=10,
            ))
            embeddings = resp.json()["embeddings"]
            if not embeddings:
                raise RuntimeError(f"candidate {candidate_id} not found")
            if embeddings[0]["ready"]:
                pending.discard(candidate_id)
        if pending:
            if time.monotonic() > deadline:
                raise TimeoutError(f"still not ready: {sorted(pending)}")
            time.sleep(2.0)


def synthesise(voice_id, text, path):
    resp = check(requests.post(
        f"{BASE_URL}/post/speech/tts",
        headers=HEADERS,
        timeout=30,
        json={"text": text, "voice_id": voice_id, "output_format": "wav", "only_audio": True},
    ))
    with open(path, "wb") as f:
        f.write(resp.content)


def convert(candidate_id, name, description=None):
    resp = check(requests.post(
        f"{BASE_URL}/voices/from-embedding",
        headers=HEADERS,
        timeout=30,
        json={"voxium_embedding_id": candidate_id, "name": name, "description": description},
    ))
    return resp.json()["uid"]


if __name__ == "__main__":
    accents = check(requests.get(f"{BASE_URL}/voice-generator/available-accents", headers=HEADERS, timeout=30)).json()["languages"]
    for language, accent, _ in TARGETS:
        assert accent in accents[language], f"{accent} is not an accent for {language}"

    print("[1/4] localizing")
    plan = {language: localize(SOURCE, language, accent)[0] for language, accent, _ in TARGETS}

    print("[2/4] waiting until ready")
    wait_until_ready(plan.values())

    print("[3/4] auditioning")
    for language, accent, line in TARGETS:
        synthesise(plan[language], line, f"candidate-{language}.wav")
        synthesise(SOURCE, line, f"source-{language}.wav")  # the same line without localization
        print(f"      candidate-{language}.wav vs source-{language}.wav  ({accent})")

    print("[4/4] listen to the files, then choose which voices to save")
    for language, accent, _ in TARGETS:
        if input(f"Save the {language} candidate as a permanent voice? [y/N] ").strip().lower() != "y":
            continue
        voice_id = convert(plan[language], f"Assistant {language.upper()}", f"Localized from {SOURCE}: {accent}")
        print(f"      {language}: voice_id={voice_id}")
    print("done")
```

## Choosing an accent

`GET /voice-generator/available-accents` returns the accents you can target,
grouped by language, with the default first. At the time of writing it covers
the five supported languages with between two and six accents each, regional
variants such as Quebecois French, Mexican Spanish or Swiss German included.
Read the list from the endpoint rather than hardcoding it, and show it to your
users as a picker: free-text accents are not accepted. Regional accents the
list does not offer, "Bristolian" or "Texan", belong to
[Voice Design](/guides/voices/voice-design), where they go in the description.

## Gender

Use `gender` to specify `male` or `female`. When you omit it:

1. A flagship voice uses its own gender tag.
2. A candidate reuses the gender its own localization was made with.
3. Otherwise no gender is specified. Clones rarely carry a gender tag,
   and converted voices do not, so send `gender` for those sources.

You can request a gender different from the source. Listen to the result to
check whether it still matches the speaker you want.

## Sources you can localize

| Source | Works | Notes |
| :- | :- | :- |
| Flagship voice | Yes | Any id from [Flagship Voices](/guides/voices/flagship-voices). |
| Instant clone | Yes, if it has a language | Clones created without a `language` return `409 The source voice has no language`. Set it once with [`PUT /voices/{voice_id}`](/guides/voices/manage-voices#update-a-voice) and retry. The language must be the one spoken in the recording: a wrong tag is accepted and quietly degrades the result. |
| Voice Design candidate or converted voice | Yes | Use the `vox_emb_` id or the converted `voice_id`. |
| Localized candidate | Yes | Localizations chain: a French candidate made from an English voice can be localized to German. The candidate must be ready (`409` otherwise). |
| Pro clone | No | `409 This voice has no embedding the voice generator can edit.` |
| Another account's voice | No | Opaque `404`. |

## Evaluating candidates

* Compare the source and candidate reading the same target-language text.
  Listen for pronunciation, accent and whether you still recognize the speaker.
* For an accent-only comparison, keep the source language and explicitly choose
  the target accent. The default accent may differ from the source's accent.
* Compare multiple candidates before saving one. Repeating a request does not
  reproduce a previous candidate.
* Use text in the candidate's target language. Keep a separate localized voice
  for each language you need.
* To clean up a noisy source rather than change its accent, use
  [Voice Enhance](/guides/voices/voice-enhance).

## Candidate lifecycle

Candidates expire after 30 days unless you save them as permanent voices with
`POST /voices/from-embedding`. Saving a candidate removes its expiry and returns
the permanent voice's `uid`. See the [conversion step](#quickstart) for the request
and response.

To remove a candidate early, use
`DELETE /voice-generator/embeddings/{embedding_id}`. A converted voice holds its
own copy. Localized candidates also keep a copy of their source, so deleting
the source after the request is accepted does not cancel generation. See the
[Voice Design candidate lifecycle](/guides/voices/voice-design#candidate-lifecycle)
for listing and cleanup details.

## Limits

| Limit | Value |
| - | - |
| Candidates per request | 1 to 5, default 1 |
| Target languages | `en`, `fr`, `es`, `pt`, `de` |
| Accents | Live list, read `GET /voice-generator/available-accents` |
| `gender` | `male` or `female` |
| Audition text | 300 characters, REST only |
| Candidate retention | 30 days, cleared once converted |
| Converted voices | Count against the custom-voice allowance |

## Errors

Errors are `{"detail": "..."}`, except `422`, which carries a list of field
errors including the valid accents for the language you asked for.

| Status | Where | Meaning |
| - | - | - |
| `401` | Any | Invalid or expired API key |
| `404` | Localize | Unknown source, or another account's voice or candidate |
| `404` | Audition | Candidate not ready yet |
| `400` | Audition | Unknown or deleted candidate, or text over 300 characters |
| `409` | Localize | Source has no language, source is a candidate that is not ready, source is a pro clone or a voice speaking an unsupported language. `detail` says which. |
| `409` | Convert | Already converted, not ready, or custom-voice allowance reached |
| `422` | Localize | Invalid `target_language`, `accent`, `gender`, `n_samples`, or an empty `src_voice` |
| `503` | Localize | Source audio could not be copied, or no voice generator is available. Retry. |

## Next steps

<CardGroup cols={2}>
  <Card title="Voice Design" icon="wand-magic-sparkles" href="/guides/voices/voice-design">
    Create the source voice from a description instead.
  </Card>

  <Card title="Custom Voices" icon="microphone" href="/guides/voices/custom-voices">
    Clone a speaker, with a language, so the clone can be localized.
  </Card>

  <Card title="Manage Voices" icon="sliders" href="/guides/voices/manage-voices">
    Set a language on an older clone, list and delete voices.
  </Card>

  <Card title="API Reference" icon="code" href="/api-reference/endpoint/localize-voice">
    The localization endpoints in full.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.