SimAWSby ShahriarLabs
Search

AWS Glue

Glue

You need to catalog data and prepare datasets for analytics without managing the full processing infrastructure.

Regional serviceConcept reference · no service console simulation

Reach for it when

  • Building ETL or ELT jobs for a data lake.
  • Maintaining a data catalog that analytics engines can use.

Do not reach for it when

  • Serving individual low-latency application requests.
  • Treating a schema crawler as a substitute for validating the meaning and quality of data.

Alternatives, and how to choose

ServicePick it instead when
AthenaQuery compatible data directly using the catalog.
EMRChoose when the workload needs more control over cluster-based data processing.

How you pay

The model
Processing capacity and duration, catalog usage and enabled features contribute to cost.
The line item that surprises people
Repeated crawls and oversized processing jobs can spend money without improving the dataset.

What trips people up

  • Partition layout and schema evolution affect downstream query behavior.
  • A catalog describes data; it does not itself guarantee transactional consistency or access to every underlying object.

Verify the live service

This page is a concept reference. Cost models are qualitative; confirm the current offering, Region and pricing before deploying.