Every provider has bad hours. Rate limits during a launch, a regional outage, a model that suddenly answers slower. If your app talks to one provider directly, its bad hour becomes yours. Routing is how you keep that from happening.
What a routing rule is
A routing rule is a small decision written once in the console: for requests that match this condition, use this upstream and this model. Your app keeps sending requests to the same endpoint, api.pecutin.co/v1, with the same key. The rule decides where they actually go.
Because the decision lives in the gateway, changing it does not require a new build of your app. You edit the rule, the gateway is notified, and the next request follows it.
Fallback: the backup path
A fallback is a second choice. If the main upstream returns an error, the request is sent to the backup instead. Your user sees an answer, not an error page, and the request is still recorded in the logs and the ledger with the upstream that actually served it.
A sensible starting point:
- A main model for the task.
- A backup from a different provider, so one provider's outage does not take out both.
- A regular look at the logs for how often fallback serves requests, so you notice when something is wrong upstream.
Routing for cost, not only for uptime
The same mechanism saves money. Many requests do not need the largest model: summaries, classification, short rewrites. Sending those to a cheaper model, and keeping the expensive one for hard tasks, is often the biggest single reduction in cost. The cost of every request is in the ledger, so you can check whether a rule actually helped.
Test the failure before it happens
A fallback that has never fired is a hope, not a plan. Turn off the main upstream on purpose in a test environment, send traffic, and confirm the backup serves it and the logs show which upstream answered.
Pertanyaan yang sering diajukan
- Do I need to change my code to use routing?
No. Your app sends requests to the same endpoint with the same key. Rules are set in the console.
- Is the fallback request charged?
Only the request that actually produces the answer is charged, at the price of the upstream that served it.