SimAWSby ShahriarLabs
Search

Key-value at any scale · Chapter 2 of 2

Why a DynamoDB scan will destroy your budget

Scans read entire tables and consume massive capacity, whereas indexes provide efficient and cost-effective access paths to your data.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

The difference between query and scan operations

Why should you avoid using the scan operation in production?

DynamoDB TableScanQueryApp ClientItem 1PK=101, Status=PendingItem 2PK=102, Status=CompletedItem 3PK=103, Status=PendingGSI (Status Index)Index PK = StatusQuerying non-key attributes in DynamoDB requires an expensive table-wide Scan operation, which is optimised to a fast, direct Query by adding a Global Secondary Index.

The client wants to fetch items with Status equal to Completed, but the table partition key is set on Item ID.

1/5
  1. 01

    How a query operates

    A query retrieves items using a specific partition key value. It targets a single partition and returns only the matching items, consuming minimal capacity and running in milliseconds.

  2. 02

    The inefficiency of a scan

    A scan reads every single item in the entire table, page by page, before filtering the results. As your table grows, a scan takes longer, costs more, and consumes your read capacity pool.

  3. 03

    Secondary indexes as alternative paths

    When you need to query data using a different attribute, you create a secondary index. The index acts as a shadow table, automatically replicated by DynamoDB with a new key structure.

Check yourself

What is the primary difference in how Query and Scan operations consume read capacity?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

Global and local secondary index mechanics

How do secondary index writes and consistency models affect table performance?

DynamoDB TableScanQueryApp ClientItem 1PK=101, Status=PendingItem 2PK=102, Status=CompletedItem 3PK=103, Status=PendingGSI (Status Index)Index PK = StatusQuerying non-key attributes in DynamoDB requires an expensive table-wide Scan operation, which is optimised to a fast, direct Query by adding a Global Secondary Index.

The client wants to fetch items with Status equal to Completed, but the table partition key is set on Item ID.

1/5
  1. 01

    Comparing GSI and LSI constraints

    Global Secondary Indexes can be created at any time and use a different partition key. Local Secondary Indexes must be created when the table is created and share the table’s partition key.

  2. 02

    Write amplification and projection cost

    When you write an item to a table, DynamoDB must write it to every secondary index. Index projections define which attributes are copied to the index; projecting unnecessary attributes increases storage costs and write amplification.

  3. 03

    The eventual consistency reality

    Global Secondary Indexes are updated asynchronously. A write to the main table is not immediately visible on the GSI, meaning reads from the index are eventually consistent.

The numbers

Maximum GSIs per table
20 indexesThe default limit for global secondary indexes per table, which is usually sufficient.
Maximum LSIs per table
5 indexesLocal secondary indexes are limited and share the 10 GB partition limit.
GSI consistency model
Eventually consistentStrongly consistent reads are not supported when querying a GSI.
LSI consistency model
Strongly or eventually consistentLSIs support strongly consistent reads because they share the partition key.

Check yourself

Why must you carefully select which attributes to project into a Global Secondary Index?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →