Find the authoritative limit
For example, ResponseCX agent creation declares
30 requests per minute in this repository’s contract. That is an endpoint-specific value,
not a platform-wide allowance. Read the relevant reference again when changing services.
The earlier generic Starter/Growth/Scale tables and universal batch-size claims have been
removed: they did not identify a service contract. If your endpoint does not publish a limit,
obtain it from your deployment administrator or support before sizing production
traffic. An omitted limit does not mean unlimited access.
Inspect the actual response
With a ResponseCX key scoped for agent reads, inspect status and headers without mutating data:200 and an agent list. On failure, keep the status,
redacted body, and request ID if present. Do not print or share your authorization header.
A service may emit Retry-After on a 429. It can contain a delay in seconds or an HTTP date.
Follow that delay if present; otherwise use bounded backoff. Do not assume every response
contains X-RateLimit-* or RateLimit-* headers, or interpret an undocumented reset field’s units.
Decide whether to retry
Batch size and capacity
Read each batch endpoint’s request schema for item limits. A limit on one operation does not apply to another. Test a representative workload on an authorized test deployment and record throughput, latency,429 frequency, and the scope shared by concurrent clients.
Bound both pending work and in-flight requests. After throttling, pausing a queue does not
re-enqueue its failed item: retry logic must explicitly make another attempt. Avoid scaling
workers simply to overcome a quota shared by the same tenant.
Continue
- Tested read retry helper — bounded attempts and
Retry-Afterhandling. - API directory — select the host and credentials.
- Errors across engines — interpret the correct error body.
- Integration test plan — verify outcomes and capacity.