Rate limits and quotas
Per-minute throttling per key, monthly allowance per team.
Two independent limits apply to every capture.
| Limit | Scope | Window | Exceeded response |
|---|---|---|---|
| Rate limit | Per API key | Rolling 60 seconds | 429 rate_limit_exceeded |
| Screenshot quota | Per team | Monthly billing period | 429 quota_exceeded |
Rate limits
Each key is throttled at its plan's per-minute allowance. Two keys on the same team each get the full allowance: throttling is per key, quota is per team.
| Plan | Requests / minute |
|---|---|
| Basic | 60 |
| Pro | 300 |
| Enterprise | 1,000 |
Every response carries the current state:
X-RateLimit-Limit: 300
X-RateLimit-Remaining: 288
X-RateLimit-Reset: 1758196800
Retry-After: 37| Header | Meaning |
|---|---|
X-RateLimit-Limit | Your plan's per-minute allowance. |
X-RateLimit-Remaining | Requests left in the current window. |
X-RateLimit-Reset | Unix timestamp when the window resets. |
Retry-After | Seconds until then. |
Cache hits and failed captures still count here, because they are requests.
Monthly quota
| Plan | Screenshots / month |
|---|---|
| Basic | 5,000 |
| Pro | 50,000 |
| Enterprise | 150,000 |
The counter covers all captures by the team, from every API key and from the dashboard's quick-generate form.
What is not counted:
- Cache hits. A capture served under
cache_ttlnever reserves a slot. - Failed captures. A render that errors releases its reserved slot.
The quota is enforced under a lock as the capture slot is reserved, so parallel requests cannot overshoot the limit.
When the period resets
The billing period is derived from your subscription's start date, not from the calendar month: a subscription started on the 12th resets on the 12th of each month. Current usage and period end are shown on your dashboard.
The 80% warning
When a capture pushes your team past 80% of its monthly allowance, the team owner is notified by email, once per billing period, so a busy month does not turn into a mailbox full of warnings.
Staying inside the limits
The highest-leverage change is caching. A cache_ttl on repeat targets
removes both quota consumption and render latency.
- Cache aggressively for pages that change slowly.
- Read
X-RateLimit-Remainingand pace your workers instead of retrying into a wall. - Honour
Retry-Afterwith exponential backoff onrate_limit_exceeded. - Use async callbacks for batch work so slow renders do not pin your own request threads.
- Treat
quota_exceededas terminal for the period: alert rather than retry.