Simulator

Tokens refill steadily; each request spends one. Empty bucket → rejected. Allows bursts up to capacity.

Token Bucketkey · user-1

Live throughput

allowed deniedrequests / second · last 40s

Benchmark

Fires 250 requests at each algorithm (20 in parallel) through real HTTP round-trips into the C++ engine, with default parameters and a fresh key.

Decision log

Every request this session — result, remaining budget and retry hints.

0 entries
1

Token Bucket

Burst-friendly throttling that still bounds your average rate.

⚙️ How it works

  1. A bucket holds up to capacity tokens and starts completely full.
  2. Tokens are added back steadily at the refill rate (tokens/second), never above capacity.
  3. Every incoming request must take one token to pass through the gate.
  4. If a token is available it is removed and the request is allowed.
  5. If the bucket is empty the request is rejected until enough time refills a token.

💡 Key insight

Because the bucket can sit full, a quiet client may spend a whole burst at once — yet over time the rate can never exceed the refill rate.

Pros

  • Allows natural bursts (great UX)
  • Bounds the long-run average
  • O(1) memory & time per key

Cons

  • A burst can briefly exceed the steady rate
  • Two knobs to tune (capacity + rate)

🎯 Best for — General-purpose API limits where the occasional short burst is perfectly fine.

🌍 In the wild — Stripe and AWS API Gateway use token-bucket style limits — e.g. “burst of 100, 10 req/s sustained.”