Lewati ke konten

Rekayasa··2 menit baca

How We Measured 2,000 Requests per Second, and What Broke First

A simulated upstream that holds each request for seconds, load from separate machines, and a money check afterwards. The method behind our capacity numbers.

Tim Pecutin

Read in English

Capacity numbers are easy to inflate. Point a load generator at a gateway whose upstream answers in one millisecond and you can report almost any figure. Here is how we measured ours, and why the method matters more than the headline.

Why the upstream has to be slow

A real model call takes seconds. That changes everything about load. At 200 requests per second with a 1 ms upstream, about 0.2 requests are alive at any moment. With a 3.75-second upstream, about 750 are. The CPU cost per request is the same; the state the gateway has to hold is three orders of magnitude larger.

So our simulated upstream waits 3 seconds plus a random 0 to 1.5 seconds before answering. Its own p95 is about 4.4 seconds, which means any latency above that is the gateway queueing, not the upstream.

Load from somewhere else

Running the load generator on the same machine as the gateway produced more wrong conclusions than anything else in this work: the two compete for CPU and the network path is artificial. Every number we publish was generated from separate machines.

The result on the production machine

On the production machine, an AMD EPYC with 8 dedicated vCPU and 32 GB RAM, load came from two separate 4-core machines.

offeredservedp95gateway overhead
1,200/s100%4,445 ms+20 ms
2,000/s100%4,685 ms+260 ms
2,400/s1,433 dropped8,925 ms+4,500 ms

The safe point is 2,000 requests per second: 360,001 requests in a row over three minutes without a single transport failure, 5xx, or 429. At that rate the machine used about 72% CPU and 8.8 GB of RAM.

The first wall was not the CPU

Before tuning, the machine stalled near 1,455 requests per second while the CPU sat at 54%. The cause was a default limit of 5,000 connections per upstream host. With requests that last 3.75 seconds, that caps throughput at about 1,333 per second. Raising that one setting removed seconds of latency.

Then we checked the money

A test that only counts errors can hide a billing bug. After the soak we counted rows: 360,001 requests produced exactly 360,001 new ledger rows, and for all 60 test accounts the balance equalled the sum of the ledger.

Pertanyaan yang sering diajukan

Is 2,000 requests per second the maximum?

It is the safe point: the highest rate with zero failures and latency close to the upstream's own. Above it, latency climbs and requests start to drop.

Do these numbers include real providers?

No. They measure the gateway itself against a simulated upstream, which is the only way to isolate its overhead.

Berhenti mengurus akun satu per satu. Mulai dari satu kunci.