SimAWSby ShahriarLabs
Search

Serverless compute · Chapter 2 of 2

Why concurrency is the true unit of Lambda capacity

Lambda scales by running multiple instances of your function in parallel, and managing this concurrency prevents database overload and application throttling.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

Understanding serverless scaling units

How does Lambda handle multiple requests arriving at the same time?

Lambda Pool (Limit: 3)App ClientsConcurrent RequestsEnvironment 1Busy (Request A)Environment 2Busy (Request B)Environment 3Busy (Request C)Throttling GateToo Many RequestsLambda scales by creating separate execution environments for concurrent requests, throttling new invocations when the account concurrency limit is reached.

Lambda scales horizontally by allocating separate execution environments to handle concurrent request streams.

1/5
  1. 01

    One container per request

    Unlike a traditional server that handles thousands of requests concurrently using threads, a single Lambda execution environment processes only one request at a time.

  2. 02

    Scaling by duplicating environments

    If ten requests arrive simultaneously, Lambda spins up ten separate containers in parallel. The number of active containers running at any moment is your function’s concurrency.

  3. 03

    The downstream database threat

    Lambda can scale from zero to thousands of concurrent environments in seconds. If these environments connect to a traditional database, they will quickly exhaust the database connection limit, crashing your database tier.

Check yourself

If a Lambda function takes 200 milliseconds to run and receives 50 requests per second, what is the average concurrency?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

Reserved and provisioned concurrency limits

How do you protect your account and application from concurrency throttling?

Lambda Pool (Limit: 3)App ClientsConcurrent RequestsEnvironment 1Busy (Request A)Environment 2Busy (Request B)Environment 3Busy (Request C)Throttling GateToo Many RequestsLambda scales by creating separate execution environments for concurrent requests, throttling new invocations when the account concurrency limit is reached.

Lambda scales horizontally by allocating separate execution environments to handle concurrent request streams.

1/5
  1. 01

    The account concurrency pool

    AWS sets a default limit on concurrent executions across all functions in an account. If one function runs out of control, it can consume the entire pool, causing all other functions in the account to fail.

  2. 02

    Reserved concurrency as a ceiling and a floor

    Reserved concurrency allocates a dedicated portion of the account limit to a function. This guarantees the function always has capacity, but also caps its scaling to prevent database overloading.

  3. 03

    Provisioned concurrency eliminates cold starts

    Provisioned concurrency pre-warms a set number of containers. When requests arrive, they run immediately in warm environments, completely eliminating cold start latency.

The numbers

Default account concurrency
1,000 concurrent executionsShared across all functions in a region. Can be increased by request.
Per-function concurrency scaling rate
1,000 environments / 10 secondsPer function, per Region, subject to account concurrency. Unused scaling capacity does not accumulate.
HTTP throttling error code
429 Too Many RequestsThe status code returned by Lambda when concurrency limits are exceeded.

Check yourself

You want to guarantee that a critical Lambda function never experiences cold starts. Which feature should you enable?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →