AWS Glue
Glue
You need to catalog data and prepare datasets for analytics without managing the full processing infrastructure.
Reach for it when
- Building ETL or ELT jobs for a data lake.
- Maintaining a data catalog that analytics engines can use.
Do not reach for it when
- Serving individual low-latency application requests.
- Treating a schema crawler as a substitute for validating the meaning and quality of data.
Alternatives, and how to choose
| Service | Pick it instead when |
|---|---|
| Athena | Query compatible data directly using the catalog. |
| EMR | Choose when the workload needs more control over cluster-based data processing. |
How you pay
- The model
- Processing capacity and duration, catalog usage and enabled features contribute to cost.
- The line item that surprises people
- Repeated crawls and oversized processing jobs can spend money without improving the dataset.
What trips people up
- Partition layout and schema evolution affect downstream query behavior.
- A catalog describes data; it does not itself guarantee transactional consistency or access to every underlying object.
Verify the live service
This page is a concept reference. Cost models are qualitative; confirm the current offering, Region and pricing before deploying.