Skip to main content
Gladia limits concurrent transcription work separately from API request rate and live session duration. Your plan determines how many jobs or sessions can run at once. If concurrency is exhausted, the live and pre-recorded APIs return HTTP 429. Use the limits below to size your workload and request a capacity increase before sustained demand exceeds your allocation. These plan limits separate usage allowance, active concurrency, and queued requests:
  • Free accounts: 3 concurrent pre-recorded jobs and 1 live session.
  • Paid defaults: 25 parallel pre-recorded jobs and 30 live sessions; up to 300 additional async requests can queue.
  • Queued requests are separate from active parallel jobs.
  • Enterprise concurrency is agreed on demand. See pricing for current plans.

Hitting your concurrency limit

When you hit your concurrent session limit, both the Live and Pre-recorded APIs return a 429 status code. HTTP 429 can mean either exhausted concurrency or API request-rate limiting. Treat them separately: reduce parallel jobs or sessions when concurrency is exhausted, and slow the request rate when request-rate limiting applies. Use bounded retries with backoff and jitter, and honor any retry header documented for the endpoint. Prefer the official SDK for Live integrations: it already handles the WebSocket lifecycle (reconnection, session continuity, retries, buffering, and related timing). See the Live quickstart for implementation details.
For pre-recorded jobs, a 200 on POST or a transcription.created webhook means the job was accepted — do not resubmit while it is still queued or processing. See Transcription process & retry policy.
Any limit above the default requires a capacity check. Contact the sales team to request an increase.
  • Usage : New accounts receive a one-time grant of €50 in credits. These credits do not renew once consumed. When they run out, top up your wallet or upgrade to a paid plan to continue.
  • Concurrency : (depending on free/paid tier) This refers to the maximum number of transcription (pre-recorder or real-time) that a user can process at the same time. For asynchronous transcriptions, Paid plan users can queue up to 300 requests, but only will still have 25 max processed concurrently.
  • Realtime session duration : (all plans) A single realtime (live) transcription session cannot exceed 3 hours. After 3 hours, the session will be terminated. When an event exceeds that duration, start a new session.
  • API level rate limit : (same for every user) Which is the number of API calls that a user can make within a particular time frame. This is to ensure that a single user or malicious actor doesn’t affect the performance of the API for all the other users.