Performance / reproducible local evidence
Show the speed.
Show the work.
A real million-product database. Bounded reads, validated prices and durable orders through Rust and PostgreSQL. Fixed arrivals include client queue time.
The million-product measurement
At 500 offered requests/s, each read workload runs for approximately 4 seconds; checkout for 2 seconds. p50, p95 and p99 include scheduling, client queue, new HTTP connections, the full response and validation. The achieved rate follows the offered rate; this is a short probe, not maximum capacity. No response cache bypasses PostgreSQL.
| Workload | p50 ms | p95 ms | p99 ms | Achieved requests/s | Errors / requests |
|---|---|---|---|---|---|
| 6-product shop / catalog page | 7.9 | 14.5 | 32.6 | 499.4 | 0 / 2000 |
| 1M products / 50-product page | 9.3 | 12.9 | 16.3 | 499.1 | 0 / 2000 |
| 1M products / product detail | 18.8 | 33.6 | 43.4 | 498.1 | 0 / 2000 |
| 1M products / localized SKU search | 16.4 | 23.7 | 28.9 | 497.5 | 0 / 2000 |
| 1M products / common-term search | 11.5 | 27.1 | 46.6 | 498.6 | 0 / 2000 |
| 1M products / 20-line cart | 13.3 | 21.4 | 30.5 | 498.7 | 0 / 2000 |
| 1M products / durable checkout | 13.2 | 17.7 | 24.2 | 496.8 | 0 / 1000 |
Apple M3 Ultra, 512 GiB host RAM; Docker engine limited to about 7.65 GiB and 24 CPUs. PostgreSQL 17 uses warm caches, fsync and synchronous commit enabled. Rust release build; Python client shares the host. A pure HTTP process and an independent outbox worker are active. Checkout changes stock and commits an order; payment is simulated.
The 100 requests/s control
The same final build also ran 1,000 requests per read workload and 500 fresh checkouts at fixed arrivals of 100/s: 6,500 measured requests, zero errors, 516 unique persisted orders including warm-up. These longer 10-second reads provide the matched follow-up to the failed cart run below.
| Workload | p50 ms | p95 ms | p99 ms | Achieved requests/s | Errors / requests |
|---|---|---|---|---|---|
| 6-product shop / catalog page | 8.3 | 12.2 | 19.2 | 100.0 | 0 / 1000 |
| 1M products / 50-product page | 9.1 | 12.4 | 21.3 | 100.0 | 0 / 1000 |
| 1M products / product detail | 11.8 | 17.5 | 25.8 | 100.0 | 0 / 1000 |
| 1M products / localized SKU search | 16.7 | 20.1 | 39.5 | 99.9 | 0 / 1000 |
| 1M products / common-term search | 10.6 | 14.7 | 18.1 | 100.0 | 0 / 1000 |
| 1M products / 20-line cart | 10.4 | 14.7 | 19.6 | 100.0 | 0 / 1000 |
| 1M products / durable checkout | 12.5 | 17.3 | 24.3 | 100.0 | 0 / 500 |
Download the complete 100/s control ↗
The failed cart run — and the fix
The first million-product run achieved only 34.9 cart reads/s at an offered 100/s. Its client admission guard rejected 634 of 1,000 scheduled reads; successful cart reads had p95 645.3 ms. PostgreSQL scanned the entire tenant for an OR/subquery. Direct indexed SKU and parent lookups removed that scan. The identical workload was then rerun: cart p95 fell to 14.7 ms with zero errors. Both tables above contain the final build.
Inspect the failed run and all client rejection records ↗
The earlier full-catalog comparison
The historical comparison below returned all 1,000 products in one response and used closed-loop clients. The current API returns at most 100 (default 50), so its smaller response is not a like-for-like full-catalog speedup. The original results and diagnostics remain available.
Historical results: all workloads
Each cell is the median of three rounds. p95 is the 95th-percentile duration inside a round, including the complete HTTP response and JSON validation. Throughput counts successful requests per elapsed second. “Clients” means concurrent closed-loop clients.
| Workload | Clients | p95 ms Before → after |
Requests / sec Before → after |
After errors / requests |
|---|---|---|---|---|
| 6-product catalog | 1 | 7.4 → 6.1 | 170.2 → 222.3 | 0 / 384 |
| 6-product catalog | 16 | 17.1 → 28.2 | 1038.7 → 753.3 | 0 / 384 |
| 6-product catalog | 64 | 83.4 → 82.4 | 812.3 → 776.6 | 0 / 384 |
| 1,000-product catalog | 1 | 94.2 → 33.9 | 11.7 → 38.8 | 0 / 384 |
| 1,000-product catalog | 16 | 233.2 → 99.6 | 115.2 → 216.0 | 0 / 384 |
| 1,000-product catalog | 64 | 715.2 → 395.0 | 106.8 → 207.4 | 0 / 384 |
| 20-line cart / 1,000 products | 1 | 82.9 → 10.0 | 13.3 → 130.5 | 0 / 384 |
| 20-line cart / 1,000 products | 16 | 177.5 → 27.5 | 119.1 → 704.2 | 0 / 384 |
| Durable checkout / 1,000 products | 16 | 25.5 → 28.2 | 671.0 → 650.7 | 0 / 192 |
What changed in the core
- Indexed translations: one lookup map replaces repeated scans of every translation. Language fallback remains field-specific; NULL and an explicitly empty string keep their different meanings.
- One product per preview: a catalog price preview searches its actual product, rather than searching the entire catalog once for every product.
- Cart-sized loading: a cart loads its own SKUs and their parents. A 20-line cart no longer hydrates 1,000 unrelated products and translations.
- Complete request intake: the catalog HTTP handler reads the POST body before returning a large response. Separate diagnostic runs exposed truncated responses when the body was left unread.
The measurement conditions
- Hardware
- Apple M3 Ultra · 512 GiB host RAM. This is a powerful development workstation.
- Software
- macOS · cargo release build with thin LTO · PostgreSQL 17 in Docker, using the project's AGE + pgvector image.
- Fixtures
- Two isolated synthetic shops: 6 and 1,000 root products. All products have German translations, identical €24.90 gross prices and ample stock. The cart contains 20 distinct items.
- Transport & client
- Local HTTP/1.1, a fresh connection per request, Python threads on the same host. Catalog measurement uses POST with no body, identically before and after. The client reads and validates every response; its overhead and shared-host contention are included.
- Warm-up
- One validated request per client before each round. Database and operating-system caches are warm; this is not a cold-start measurement.
- Checkout
- Every measured checkout has a fresh cart and unique idempotency key. Stock is updated and the order is committed in PostgreSQL. Payment is simulated; no payment-provider network request is timed. All measured order IDs are checked directly in the database.
Failed attempts are part of the record
Initial catalog runs sending a JSON POST body timed out or received connection resets. They were aborted, not counted as clean benchmark runs. Body-free control requests completed; the HTTP body-consumption fix was then tested with the original JSON-body case. Read the diagnostic record and follow-up result.
Run it yourself
The reproducible setup documents the isolated database, build commands and benchmark script. Raw files include every successful latency, failures, source revision/diff and binary hashes.
What these numbers do not establish
Production capacity, internet latency, cold-cache performance, a many-shop fleet, distributed deployment, real payment-provider performance and local LLM generation speed remain unmeasured here. There is no Shopware speed comparison. These are bounded local measurements, not a production SLA or maximum server capacity.