The most common way an AI bill gets out of hand is not a price change. It is one key used everywhere: in the backend, in a staging server, in a script on someone's laptop. When that key leaks or a loop goes wrong, nothing stops it.
One key per job
Pecutin keys look like sk-pc-…, and you can create as many as you need in the console. A good split:
- One key per service, so a chatbot and a summarizer do not share a limit.
- Separate keys for staging and production.
- One key per customer if you resell access, so each customer's usage is visible on its own.
Quotas: a ceiling on spend
A quota caps how much a key can spend. When a key reaches its quota, the next request is refused with a clear reason. The other keys are not affected, so one runaway service cannot drain the balance that production depends on.
Quotas are also a planning tool. If a new feature has a budget, put that budget on its key and you will know the moment it is used up.
Rate limits: a ceiling on speed
A rate limit caps how many requests a key can send in a period. It protects you from a retry loop that sends thousands of requests in a minute, and it protects an expensive model from being flooded by one client.
When a key goes over its rate limit, the request is refused with a 429 and a reason, so your code can back off instead of guessing.
Rotate without a release
Because your app only knows its own key, replacing a leaked key is quick: create a new one, put it in the service's configuration, and revoke the old one. No other service has to change.
Pertanyaan yang sering diajukan
- What does my app receive when a limit is reached?
An error response with a clear reason, so you can tell a quota limit from a rate limit and handle each one.
- Can I see which key spent what?
Yes. Usage and cost are shown per key in the console, and every request is in the ledger.