> ## Documentation Index
> Fetch the complete documentation index at: https://gladia-95-fix-geo-concurrency.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# API rate limits and transcription concurrency

> Check Gladia transcription concurrency, queue capacity and live session duration. Understand 429 responses and how to request more capacity.

Gladia limits concurrent transcription work separately from API request rate and live session duration. Your plan determines how many jobs or sessions can run at once. If concurrency is exhausted, the live and pre-recorded APIs return HTTP 429. Use the limits below to size your workload and request a capacity increase before sustained demand exceeds your allocation.

| Plan type | Usage | Max Transcriptions in concurrency <br />(pre-recorded) | Max Transcriptions in concurrency (Live) |
| - | - | - | - |
| **Enterprise** | Unlimited | On demand | On demand |
| **Paid** | Unlimited | 25 | 30 |
| **Free** | €50 credits (one-time) | 3 | 1 |

These plan limits separate **usage allowance**, **active concurrency**, and **queued requests**:

* **Free accounts:** 3 concurrent pre-recorded jobs and 1 live session.
* **Paid defaults:** 25 parallel pre-recorded jobs and 30 live sessions; up to 300 additional async requests can queue.
* Queued requests are separate from active parallel jobs.
* **Enterprise** concurrency is agreed on demand. See [pricing](https://www.gladia.io/pricing) for current plans.

## Hitting your concurrency limit

When you hit your concurrent session limit, both the **Live** and **Pre-recorded** APIs return a **429** status code.

HTTP 429 can mean either exhausted concurrency or API request-rate limiting. Treat them separately: reduce parallel jobs or sessions when concurrency is exhausted, and slow the request rate when request-rate limiting applies. Use bounded retries with backoff and jitter, and honor any retry header documented for the endpoint.

Prefer the [official SDK](/chapters/how-to-use-gladia/sdk) for Live integrations: it already handles the WebSocket lifecycle (reconnection, session continuity, retries, buffering, and related timing). See the [Live quickstart](/chapters/live-stt/quickstart) for implementation details.

<Note>
  For pre-recorded jobs, a **200** on `POST` or a [`transcription.created`](/api-reference/v2/pre-recorded/webhook/created) webhook means the job was accepted — do not resubmit while it is still queued or processing. See [Transcription process & retry policy](/chapters/pre-recorded-stt/transcription-process).
</Note>

### Paid-plan concurrency defaults

| API | Default |
| - | - |
| Live (WebSocket) | 30 concurrent sessions |
| Pre-recorded (async) | 25 parallel jobs + 300 queued |

<Note>
  Any limit above the default requires a capacity check. [Contact the sales team](https://www.gladia.io/contact) to request an increase.
</Note>

* **Usage** :

  New accounts receive a **one-time grant of €50 in credits**. These credits do **not** renew once consumed. When they run out, top up your wallet or upgrade to a paid plan to continue.

* **Concurrency** : (depending on free/paid tier)

  This refers to the maximum number of transcription (pre-recorder or real-time) that a user can process at the same time.
  For asynchronous transcriptions, Paid plan users can queue up to 300 requests, but only will still have 25 max processed concurrently.

* **Realtime session duration** : (all plans)

  A single realtime (live) transcription session cannot exceed **3 hours**. After 3 hours, the session will be terminated. When an event exceeds that duration, start a new session.

* **API level rate limit** : (same for every user)

  Which is the number of API calls that a user can make within a particular time frame.
  This is to ensure that a single user or malicious actor doesn't affect the performance of the API for all the other users.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.