Skip to content

Scalability, Latency & Throughput

Core distinction

  • Latency: How long one operation takes.
  • Throughput: How much work the system processes per unit time.
  • Concurrency: How many operations are in progress at once.
  • Scalability: How system capacity changes as demand changes.

Architect's questions

  • What is average traffic?
  • What is peak traffic?
  • What is the expected growth?
  • What is the latency target?
  • Which operations are CPU-bound?
  • Which are I/O-bound?
  • Where is the bottleneck?
  • Can the bottleneck scale horizontally?

Capacity estimation

Capacity estimates connect demand to resource requirements:

Users
  ↓
Requests / user / day
  ↓
Requests / day
  ↓
Average RPS
  ↓
Peak RPS
  ↓
Compute / storage / network requirements

Design scenario

Design an API that must support:

  • 10 million users
  • 1 million daily active users
  • 20 requests per active user per day
  • 10× peak traffic compared with the daily average

Calculate approximate average and peak requests per second and identify likely bottlenecks.

Further topics

  • Little's Law
  • Horizontal vs vertical scaling
  • Load balancing
  • Queuing
  • Backpressure
  • Capacity planning