Rate limits

What is throttled, at what rate, and how to behave when you hit it.

Public surfaces are rate-limited per source address. Exceeding a limit returns 429 with a Retry-After header saying how many seconds to wait.

HTTP/1.1 429 Too Many Requests
Retry-After: 23

Default limits

These are defaults. A self-hosted deployment can change them, so read Retry-After rather than hard-coding the numbers below.

SurfaceLimit
Auth routes — signup, login, password reset, magic link, email verification, OAuth exchange30 requests / 60s per IP
Public widget chat (/v1/public/agents/{key}/chat)60 requests / 60s per IP
Channel webhooks (Telegram, Meta, Slack, Discord)120 requests / 60s per source
n8n callback (/v1/tools/n8n/callback)120 requests / 60s per source
Invitation preview30 requests / 60s per IP

Authenticated API routes are not throttled per request in the same way. They are bounded by your organization's plan limits instead — message counts, agent counts, and which features are available.

Plan limits

Separate from rate limiting, and they fail differently. When an organization exhausts its plan's message allowance, the agent stops replying rather than erroring: the conversation is routed to the human inbox with a badge explaining why.

That choice is deliberate. A visitor should never see an API error because you hit a quota — they should get a person.

Calls that would exceed a plan limit return 402 with a plan_limit code.

Behaving well

Read Retry-After. It is the actual answer. Retrying immediately after a 429 makes the problem worse and, on a shared IP, makes it worse for someone else.

Back off exponentially on 429 and 5xx, with jitter:

async function withRetry(fn, attempts = 5) {
  for (let i = 0; i < attempts; i++) {
    const res = await fn();
    if (res.status !== 429 && res.status < 500) return res;
    const retryAfter = Number(res.headers.get("Retry-After"));
    const wait = retryAfter
      ? retryAfter * 1000
      : Math.min(2 ** i * 500, 30_000) * (0.5 + Math.random());
    await new Promise((r) => setTimeout(r, wait));
  }
  throw new Error("giving up after repeated rate limiting");
}

Never retry a 4xx other than 429. Nothing about the request will have changed.

Behind a proxy

Limits are per client IP. If your requests reach Vicero through a proxy, the deployment must be configured to trust that proxy's forwarded headers — otherwise every visitor counts as one client and a single busy customer exhausts the limit for everyone.

This matters most for the widget, where the "client" is a real visitor's browser. If you self-host, see Self-hosting.