> ## Documentation Index
> Fetch the complete documentation index at: https://docs.jiekou.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# ElevenLabs Speech to Text V1

Transcribe audio or video files. When use\_multi\_channel is true and the uploaded audio has multiple channels, a 'transcripts' object is returned, with one transcript per channel. Otherwise, a single transcription result is returned.

## Request Headers

<ParamField header="Content-Type" type="string" required={true}>
  Enum value: `application/json`
</ParamField>

<ParamField header="Authorization" type="string" required={true}>
  Bearer authentication format: Bearer \{\{API key}}.
</ParamField>

## Request Body

<ParamField body="seed" type="integer" nullable={true}>
  If specified, the system will make a best effort to sample deterministically. Requests with the same seed and parameters should return the same result, but absolute determinism is not guaranteed. Must be an integer between 0 and 2147483647.

  Value range: \[0, 2147483647]
</ParamField>

<ParamField body="diarize" type="boolean" default={false}>
  Whether to annotate the current speaker in the uploaded file.
</ParamField>

<ParamField body="file_format" type="string" default="other">
  The input audio format. Options are 'pcm\_s16le\_16' or 'other'. pcm\_s16le\_16 requires the audio to be 16kHz sample rate, 16-bit integer, mono, little-endian format, and has lower latency compared to encoded waveforms.

  Allowed values: `pcm_s16le_16`, `other`
</ParamField>

<ParamField body="temperature" type="number" nullable={true}>
  Controls the randomness of the transcription output. The value ranges from 0.0 to 2.0; higher values make results more diverse and less deterministic. If omitted, the selected model's default temperature is used (usually 0).

  Value range: \[0, 2]
</ParamField>

<ParamField body="num_speakers" type="integer" nullable={true}>
  The maximum number of speakers in the uploaded file. This can be used to help distinguish speakers, with support for up to 32 speakers.

  Value range: \[1, 32]
</ParamField>

<ParamField body="language_code" type="string" nullable={true}>
  Specify the ISO-639-1 or ISO-639-3 language code of the audio file. Providing this in advance can sometimes improve transcription performance. The default is null, which automatically detects the language.
</ParamField>

<ParamField body="tag_audio_events" type="boolean" default={true}>
  Whether to tag audio events such as (laughter) and (footsteps) in the transcription.
</ParamField>

<ParamField body="cloud_storage_url" type="string" required={true} nullable={true}>
  The HTTPS link to the file to be transcribed. Exactly one of file and cloud\_storage\_url must be provided. The file must be accessible over HTTPS and smaller than 2GB. Any valid HTTPS address is supported, including cloud storage (AWS S3, GCS, Cloudflare R2, etc.), CDNs, or other HTTPS sources. Presigned links with tokens or authentication via URL query parameters are supported.
</ParamField>

<ParamField body="use_multi_channel" type="boolean" default={false}>
  Whether the audio file is multi-channel and each channel contains only a single speaker. When enabled, each channel is transcribed independently and the results are combined. Each word in the output content contains a channel\_index field. Up to 5 channels are supported.
</ParamField>

<ParamField body="diarization_threshold" type="number" nullable={true}>
  The diarization threshold. A higher value lowers the probability that one person is split into multiple speakers, but increases the probability that different people are merged into one speaker (fewer speakers detected). A lower value increases the probability that one person is split into multiple speakers, but reduces the probability that different people are merged into one speaker (more speakers detected). Can only be set when diarize=True and num\_speakers=None. The default is None, and the threshold is selected based on the model id (usually 0.22).

  Value range: \[0.1, 0.4]
</ParamField>

<ParamField body="timestamps_granularity" type="string" default="word">
  The granularity of timestamps in the transcription. 'word' provides word-level timestamps, while 'character' provides timestamps for each character.

  Allowed values: `none`, `word`, `character`
</ParamField>

## Response Information

<Note>
  The response may be one of the following response types:
</Note>

<Accordion title="Response Type 1">
  <ResponseField name="text" type="string" required={true}>
    The raw transcribed text.
  </ResponseField>

  <ResponseField name="words" type="object[]" required={true}>
    A list of words and their timing information.

    <Expandable title="properties" defaultOpen={true}>
      <ResponseField name="end" type="number" required={false}>
        The end time of the word or sound in the audio, in seconds.
      </ResponseField>

      <ResponseField name="text" type="string" required={true}>
        The transcribed word or sound content.
      </ResponseField>

      <ResponseField name="type" type="string" required={true}>
        The type of this word or sound. 'audio\_event' is used for non-word sounds, such as laughter or footsteps.

        Allowed values: `word`, `spacing`, `audio_event`
      </ResponseField>

      <ResponseField name="start" type="number" required={false}>
        The start time of the word or sound in the audio, in seconds.
      </ResponseField>

      <ResponseField name="logprob" type="number" required={true}>
        The log probability when predicting this word. The logprob range is \[-infinity, 0]; higher values indicate that the model is more confident in its prediction.
      </ResponseField>

      <ResponseField name="characters" type="object[]" required={false}>
        The characters that make up the word and their corresponding timing information.

        <Expandable title="properties" defaultOpen={true}>
          <ResponseField name="end" type="number" required={false}>
            The end time of the character in the audio, in seconds.
          </ResponseField>

          <ResponseField name="text" type="string" required={true}>
            The transcribed character content.
          </ResponseField>

          <ResponseField name="start" type="number" required={false}>
            The start time of the character in the audio, in seconds.
          </ResponseField>
        </Expandable>
      </ResponseField>

      <ResponseField name="speaker_id" type="string" required={false}>
        The unique identifier of the speaker corresponding to this word.
      </ResponseField>
    </Expandable>
  </ResponseField>

  <ResponseField name="channel_index" type="integer" required={false}>
    The channel index corresponding to this transcript (valid for multi-channel audio).
  </ResponseField>

  <ResponseField name="language_code" type="string" required={true}>
    The detected language code (for example, 'eng' for English).
  </ResponseField>

  <ResponseField name="transcription_id" type="string" required={false}>
    The unique transcription ID for this response.
  </ResponseField>

  <ResponseField name="language_probability" type="number" required={true}>
    The confidence of language detection (between 0 and 1).
  </ResponseField>
</Accordion>

<Accordion title="Response Type 2">
  <ResponseField name="transcripts" type="object[]" required={true}>
    A list of transcripts corresponding to each audio channel. Each transcript contains the text for its channel and word-level details.

    <Expandable title="properties" defaultOpen={true}>
      <ResponseField name="text" type="string" required={true}>
        The raw transcribed text.
      </ResponseField>

      <ResponseField name="words" type="object[]" required={true}>
        A list of words and their timing information.

        <Expandable title="properties" defaultOpen={true}>
          <ResponseField name="end" type="number" required={false}>
            The end time of the word or sound in the audio, in seconds.
          </ResponseField>

          <ResponseField name="text" type="string" required={true}>
            The transcribed word or sound content.
          </ResponseField>

          <ResponseField name="type" type="string" required={true}>
            The type of this word or sound. 'audio\_event' is used for non-word sounds, such as laughter or footsteps.

            Allowed values: `word`, `spacing`, `audio_event`
          </ResponseField>

          <ResponseField name="start" type="number" required={false}>
            The start time of the word or sound in the audio, in seconds.
          </ResponseField>

          <ResponseField name="logprob" type="number" required={true}>
            The log probability when predicting this word. The logprob range is \[-infinity, 0]; higher values indicate that the model is more confident in its prediction.
          </ResponseField>

          <ResponseField name="characters" type="object[]" required={false}>
            The characters that make up the word and their corresponding timing information.

            <Expandable title="properties" defaultOpen={true}>
              <ResponseField name="end" type="number" required={false}>
                The end time of the character in the audio, in seconds.
              </ResponseField>

              <ResponseField name="text" type="string" required={true}>
                The transcribed character content.
              </ResponseField>

              <ResponseField name="start" type="number" required={false}>
                The start time of the character in the audio, in seconds.
              </ResponseField>
            </Expandable>
          </ResponseField>

          <ResponseField name="speaker_id" type="string" required={false}>
            The unique identifier of the speaker corresponding to this word.
          </ResponseField>
        </Expandable>
      </ResponseField>

      <ResponseField name="channel_index" type="integer" required={false}>
        The channel index corresponding to this transcript (valid for multi-channel audio).
      </ResponseField>

      <ResponseField name="language_code" type="string" required={true}>
        The detected language code (for example, 'eng' for English).
      </ResponseField>

      <ResponseField name="transcription_id" type="string" required={false}>
        The unique transcription ID for this response.
      </ResponseField>

      <ResponseField name="language_probability" type="number" required={true}>
        The confidence of language detection (between 0 and 1).
      </ResponseField>
    </Expandable>
  </ResponseField>

  <ResponseField name="transcription_id" type="string" required={false}>
    The unique transcription ID for this response.
  </ResponseField>
</Accordion>
