Sale opens at 8:00 PM. By 8:03 the homepage is still spinning.
An apparel brand puts 300 limited units on sale at 8:00 PM. At 8:00:12 the support LINE account fills with "it just keeps spinning." At 8:01 database connections hit the ceiling, at 8:03 the homepage stops loading, and a restart at 8:11 brings it back. Final tally: 47 of 300 units sold, 12 oversold orders refunded by phone one at a time, and NT$180,000 of ad spend traded for a week of complaints. The verdict is always "the server was too small." Eight times out of ten, it was not.
Who needs this, and who does not
Good fit: limited-stock drops (thousands of people hitting one SKU within five minutes), course and event registration (30-50x a normal day at the moment of opening), tickets and appointments, where finite resources always produce row contention.
Not a fit: under 200 orders a day with no identifiable peak, or sites fully hosted on SHOPLINE or Shopify, where the platform absorbs the peak. There, money spent on conversion returns far more.
Four approaches, each treating a different layer
| Approach | Layer it fixes | Cost band (estimate) | Limits |
|---|---|---|---|
| Scale up the server | CPU and memory only; no effect on row locks | +NT$5,000–30,000 per month | The database is still a single point |
| SaaS commerce platform (SHOPLINE / Shopify / 91APP) | All three layers managed, including CDN and stock decrement | From roughly NT$60,000 a year plus revenue share | Limited customization, hard to keep ERP logic |
| Cloudflare Waiting Room only | Front end and part of the app layer | Tens to a few hundred US dollars a month | Checkout still jams after admission |
| Full in-house capacity programme | All three layers | NT$260,000–650,000 one-off | Needs 6-9 weeks and real staging |
A bigger server cannot fix a row lock, and a waiting room cannot fix overselling.
Five phases, from audit to dress rehearsal
Phase 1. Audit and bottleneck location (1 week) Separate the three ways a site dies: CDN buried by static assets, worker pools exhausting their queue, connections and row locks deadlocking. Deliverables: architecture diagram, bottleneck list, slow query report (Laravel Telescope).
Phase 2. Baseline load testing (1-2 weeks) Concurrency slots = peak RPS x average response time (seconds): 3,000 shoppers each sending one request within five seconds is a 600 RPS peak, and at 200ms you need roughly 120 slots. Use k6, JMeter or Locust against the peak curve, on production-like staging, exercising full checkout. Deliverable: capacity report.
Phase 3. Cache, queue and locks (2-4 weeks) Layer the cache: CDN edge, full page, fragment, Redis, OPcache. Use Cache-Control with stale-while-revalidate to keep product pages cacheable while stock counts come from a separate API call. Move email, invoicing and push into Laravel Queue, leaving only stock decrement and order creation on the checkout path. Deliverable: locking strategy document.
Phase 4. Waiting room and degradation (1 week) Adopt Cloudflare Waiting Room, Queue-it or a self-built token bucket, and give payment, SMS and logistics APIs timeouts, retry caps and circuit breakers. Deliverable: degradation runbook.
Phase 5. Dress rehearsal (1 week) Seven days out, run the full flow including ECPay sandbox payment and refund, and kill one node deliberately to confirm degradation engages. Deliverable: on-call runbook.
Real cost breakdown (estimates, NTD)
- Load testing and capacity report: NT$60,000–150,000 (2-3 rounds)
- Cache and queue rework: NT$120,000–300,000
- Waiting room: Cloudflare plans run tens to a few hundred US dollars a month; self-built NT$80,000–150,000
- Oversell prevention: NT$80,000–200,000
- On-call standby on sale day: NT$15,000–40,000 per day
Hidden costs
- A production-like staging replica at NT$3,000–15,000 per month
- Peak-hour cloud spend can be 3-8x a normal day, and belongs in the campaign margin
- High-concurrency gateway plans are negotiated separately
- SMS and push unit prices are low, but totals overshoot
What clients expect vs what actually happens
- "A bigger server solves it." What falls over is the connection ceiling and the row lock on one stock record; 8 cores to 32 delays the jam by 40 seconds.
- "Auto-scaling will save us." New instances need 2-5 minutes to warm up and cannot rescue a 30-second spike; pre-scaling plus queueing works.
- "Going down is the worst outcome." Overselling is worse: refunds, vouchers and apologies cost more than the orders you missed.
Five traps and how to avoid them
- Stock counts get cached: the page says "3 left" long after it hit zero. Fix: fetch stock from a separate API with a 5-10 second TTL.
- Pessimistic locks that are too broad: an over-wide SELECT ... FOR UPDATE queues everyone at 600 RPS. Fix: lock the single SKU row, or use optimistic locking with a version column and retries.
- Redis decrement without reconciliation: atomic DECR handles volume best, but a failed database write leaves numbers out of sync. Fix: add a reconciliation job and compensating transactions.
- Third parties with no timeout: a gateway that takes 10 seconds eats the whole worker pool. Fix: 3-5 second timeouts, retry caps, circuit breakers, and a "create the order now, settle payment later" mode.
- Launching without observability: when it breaks you are guessing. Fix: APM plus a dashboard for the four golden signals (latency, traffic, errors, saturation) before launch.
Success metrics and a 90-day roadmap
Day 30: load test complete, p95 under 800ms at peak, error rate under 1%, LCP still holding 2.5 seconds. Day 60: queue backlog drains within 3 minutes after the peak, zero oversold orders, admission rate within 15% of measured capacity. Day 90: one real campaign completed, with cloud cost per order compared against the baseline.
Decision checklist
- ☐ I know last campaign's peak RPS and p95
- ☐ We have staging with production-like data volume
- ☐ Load tests cover checkout and payment callbacks
- ☐ I know our database connection ceiling
- ☐ Stock decrement has an explicit locking strategy
- ☐ Email, invoicing and push are off the checkout path
- ☐ Product page and stock caching are handled separately
- ☐ Every third-party API has a timeout and a retry cap
- ☐ A degradation mode exists and someone can trigger it
- ☐ APM is installed, a rehearsal is set for 7 days out, on-call confirmed
FAQ
Can a waiting room alone carry us through?
A waiting room only controls the rate people come in. If checkout still takes 3 seconds and stock decrement still locks a wide range of rows, the jam has moved from the door to the counter, and overselling continues.
How large does the load test need to be?
Three scenarios at minimum: expected peak, twice expected peak, and the vertical 30-second curve at the moment the sale opens. The damage lands in the first 60 seconds, so averages tell you nothing.
Pessimistic locking, optimistic locking or Redis atomic decrement?
Below 100 RPS, pessimistic locking is simplest. Between 100 and 500 RPS, optimistic locking with retries keeps conflict rates manageable. Above 500 RPS, use Redis DECR to absorb volume and persist orders asynchronously, but reconciliation afterwards is mandatory.
Before your next drop, get a capacity check
ScriptWalker offers a Campaign Capacity Check from NT$60,000, covering load testing, bottleneck location and a capacity report; High-Concurrency Architecture Rework starts at NT$180,000 and covers cache layering, queues, locking strategy and waiting room setup. A full rework takes 6-9 weeks, so start 10-12 weeks out.
- Email:[email protected]
- 電話:0916-224-047
- LINE:@ufv9089p