Skip to content
Journal / education

Rate limits and 429s: strategies for AI agents in 2026

Why 429 errors deserve different handling than 500s, and when AI agents should back off, fail over, or spread load across multiple API providers.

A 429 is the most misunderstood error code in agent infrastructure. Teams treat it like a 500 (retry with backoff, hope it clears), but a 429 is not the provider being broken. It is the provider being healthy and telling you, specifically you, to slow down. That distinction changes what the correct response is, and most agent codebases get it wrong in one of two directions: they retry too aggressively and make it worse, or they back off too politely and stall a pipeline that had other options.

I run a tool router, which means I watch 429s across roughly 30 providers all day. Here is how I think about them.

Why agents hit rate limits harder than apps do

A human-facing app generates requests at human speed. An agent generates requests at loop speed. One agent run can fire dozens of search, scrape, and parse calls in a few seconds, and a fleet of concurrent agent sessions multiplies that. Provider rate limits were mostly designed around the first pattern, so agents slam into them in bursts: everything works in testing, then a batch job launches 50 sessions and every one of them starts eating 429s at once.

Worse, naive retry logic synchronizes the pain. Fifty sessions get limited at the same moment, all sleep the same fixed interval, and all retry at the same moment. Congratulations, you have built a self-inflicted thundering herd.

The three strategies

Backoff: correct, but only when you have time

Exponential backoff with jitter is the textbook answer and it works, with two caveats. First, the jitter is not optional; it is the part that breaks retry synchronization. Second, backoff assumes the request can wait. For an offline indexing job, waiting is free. For an agent mid-task, or worse a voice agent mid-conversation, every second of backoff is latency the user experiences. Backoff is the right tool when the work is patient and the provider is the only source of that capability.

Failover: the underused option

Here is the thing about tool APIs that makes them different from your own database: most capabilities have three or more substitutable providers. If Serper rate-limits you, Brave and Perplexity can answer the same search. A 429 from one vendor does not mean the query cannot be answered right now; it means this vendor cannot answer it right now. Failing over to a second provider turns a wait into a slightly different (often slightly pricier) success, and for latency-sensitive agents that trade is almost always worth it. I wrote about the broader case in why AI agents need failover; rate limits are the most common trigger in practice, because unlike outages they happen when everything is working as designed.

The catch is that failover on 429s needs discipline. If provider A limits you and you dump 100% of traffic on provider B instantly, you may just trip B’s limits too. Which brings me to the third strategy.

Spreading: don’t hit the limit in the first place

If your steady-state volume is anywhere near a single provider’s ceiling, the durable fix is to spread load across providers proactively rather than reactively. Two providers at 50% of their limits beat one provider at 100% in every dimension that matters: you have headroom for bursts, an always-warm fallback path, and real data on both vendors’ behavior instead of a cold standby you have never tested. The cost is maintaining two integrations, which is exactly the tax a normalized routing layer exists to remove.

Why 429s and 500s need different plumbing

This is the part I would tattoo on every agent codebase. A 500 is evidence the provider is unhealthy; enough of them and you should stop sending traffic entirely. Our router’s circuit breaker works this way: a provider exceeding a 30% error rate over a five minute window gets skipped, because continuing to send requests to a struggling service helps nobody.

A 429 is evidence of the opposite. The provider is fine; your allocation is exhausted. Feeding 429s into the same circuit breaker as 500s would mark healthy providers as dead exactly when you need routing to be smart. The correct model treats them separately: 500s and timeouts degrade a provider’s health score, while 429s trigger immediate failover for that request without poisoning the provider’s standing. It stays in the rotation, because the very next request (different key, different minute, different quota window) may sail through. I go deeper on the health-tracking side in circuit breakers for AI agents.

What this looks like when it’s someone else’s problem

For what it’s worth, this whole taxonomy is baked into how route.tools handles a request: a rate-limited provider triggers automatic failover, up to three providers per call, while the circuit breaker only reacts to genuine failures. Each response includes the full routing.attempted chain, so when a 429 did occur you can see it, see who caught the request instead, and see what it cost. Per-category wall-clock budgets cap the total time failover is allowed to spend, so a bad afternoon degrades gracefully instead of hanging your agent.

You can build all of this yourself, and if you only depend on one or two providers you probably should start with plain backoff plus jitter and nothing else. The complexity only pays for itself once multiple substitutable providers exist for your workload, and at that point the routing docs describe the behavior you would otherwise be writing by hand.

If you want to see which categories have enough substitutable providers to make failover realistic, the search API comparison is a good place to start.