Amazon Athena
Athena
You need to run standard SQL queries directly against files stored in S3 without loading them into a database.
Reach for it when
- Querying VPC flow logs or CloudTrail events stored in S3 to investigate a security incident.
- Analyzing ad-hoc CSV, JSON, or Parquet datasets generated by background application exports.
- Building quick dashboards on top of data lakes using BI tools like QuickSight.
Do not reach for it when
- Serving low-latency, high-frequency queries to end users of a web application — use RDS or DynamoDB instead.
- Running complex transactional queries requiring frequent insert, update, or delete operations — use RDS instead.
- Analyzing massive, petabyte-scale data warehouses with complex, multi-stage ETL pipelines — use Redshift instead.
Alternatives, and how to choose
| Service | Pick it instead when |
|---|---|
| Redshift | Choose it when you need a dedicated, high-performance data warehouse for continuous complex analytics. |
| RDS | Choose it when running standard transactional applications that require sub-second query speeds. |
How you pay
- The model
- Pay per TB of data scanned by your SQL queries, rounded up to the nearest megabyte.
- The line item that surprises people
- Running unoptimized queries against uncompressed raw JSON files in S3 can result in high scanning costs.
What trips people up
- Failing to partition your data in S3 results in Athena scanning the entire bucket for every query, inflating costs.
- Athena does not support standard indexes; query speed depends entirely on file size, compression, and partitioning.
- Simultaneous query execution limits are regional; submitting too many parallel queries will trigger throttling errors.
Verify the live service
This page is a concept reference. Cost models are qualitative; confirm the current offering, Region and pricing before deploying.