ShahriarLabs Simulators · SimAWS
Find your next explanation
Search the learning library by concept, service or problem.
126 lessons, paths, guides and experiments. Search by topic or problem.
- Service reference
IAM
Securely manage access to AWS services and resources. users roles policies permissions mfa identity
Explore - Service reference
S3
Scalable storage in the cloud. object storage bucket static website glacier storage
Explore - Service reference
VPC
Isolated cloud resources. network subnet route table gateway private network cidr
Explore - Service reference
EC2
Virtual servers in the cloud. virtual machine server compute instance ami
Explore - Service reference
EBS
Block storage volumes for EC2 instances. block storage volume disk snapshot san
Explore - Service reference
ELB
Distribute incoming traffic across multiple targets. load balancer alb nlb traffic distribution routing
Explore - Service reference
EC2 Auto Scaling
Scale EC2 capacity up or down automatically. auto scaling scale group capacity elasticity launch template
Explore - Service reference
Lambda
Run code without thinking about servers. serverless function faas microservices run code
Explore - Service reference
API Gateway
Build, deploy, and manage APIs at any scale. api rest http websocket endpoint serverless
Explore - Service reference
DynamoDB
Managed NoSQL database. nosql key-value document database serverless tables
Explore - Service reference
SQS
Managed message queues for decoupling applications. queue message queue decoupling fifo pubsub
Explore - Service reference
SNS
Pub/sub messaging, SMS, email, and mobile push notifications. pubsub notification push notification sms topic
Explore - Service reference
EventBridge
Serverless event bus for SaaS and AWS services. event bus serverless event-driven schema registry rules
Explore - Service reference
Step Functions
Visual workflows for serverless applications. workflow state machine orchestration serverless saga
Explore - Service reference
CloudWatch
Monitor resources and applications. monitoring metrics logs alarms dashboard
Explore - Service reference
CloudTrail
Track user activity and API usage. audit compliance api logging activity tracker governance
Explore - Service reference
Config
Record and evaluate configurations of AWS resources. compliance configuration history audit resource inventory
Explore - Service reference
Systems Manager
Configure and manage EC2 and on-premises resources. parameter store run command patch manager ops center session manager
Explore - Service reference
Billing and Cost Management
Monitor your AWS cost and usage. cost explorer invoice budgets pricing payment
Explore - Service reference
RDS
Managed relational database service. relational database sql mysql postgresql oracle sql server
Explore - Service reference
Aurora
High performance managed relational database. relational database mysql postgresql serverless database
Explore - Service reference
CloudFront
Global content delivery network. cdn content delivery cache edge location latency
Explore - Service reference
Route 53
Scalable Domain Name System. dns domain registration routing health checks name server
Explore - Service reference
KMS
Create and control keys used to encrypt your data. encryption keys cryptography security secrets
Explore - Service reference
Secrets Manager
Rotate, manage, and retrieve database credentials and API keys. secrets credentials password manager rotation security
Explore - Service reference
Certificate Manager
Provision, manage, and deploy SSL/TLS certificates. ssl tls certificates https security
Explore - Service reference
ECS
Run containerized applications. docker containers fargate microservices orchestration
Explore - Service reference
ECR
Easily store, manage, and deploy container images. docker registry images container registry repository
Explore - Service reference
EKS
Managed Kubernetes service. kubernetes k8s containers orchestration fargate
Explore - Service reference
CloudFormation
Model and provision AWS resources using Infrastructure as Code. infrastructure as code iac templates stack automation
Explore - Service reference
Organizations
Centralized management for multiple AWS accounts. multi-account consolidated billing ou scp governance
Explore - Service reference
WAF
Protect web applications from common web exploits. firewall web application firewall security ddos acl rules
Explore - Service reference
EFS
Serverless file storage for EC2 and containers. file storage nfs network file system shared storage
Explore - Service reference
ElastiCache
In-memory data store and cache. cache redis memcached in-memory database performance
Explore - Service reference
Athena
Query data in S3 using SQL. sql query serverless query s3 query interactive analytics
Explore - Service reference
Glue
Simple, scalable, and serverless data integration. etl data catalog data integration serverless spark schema registry
Explore - Service reference
Redshift
Fast, simple, cost-effective data warehousing. data warehouse columnar database olap analytics big data
Explore - Service reference
Kinesis
Process real-time streaming data at scale. streaming data real-time event stream data ingestion
Explore - Service reference
Cognito
Identity management for your apps. user pool authentication identity provider oauth federation
Explore - Service reference
CodePipeline
Continuous delivery service for fast and reliable updates. cicd pipeline deployment continuous integration release
Explore - Service reference
CodeBuild
Build and test code with scalable pay-as-you-go servers. compiler build server test code continuous integration docker build
Explore - Service reference
CodeDeploy
Automate code deployments to maintain application uptime. deployment automate deploy rolling updates blue green
Explore - Service reference
X-Ray
Analyze and debug production, distributed applications. tracing debugging performance analysis microservices
Explore - Service reference
SageMaker
Build, train, and deploy machine learning models. machine learning deep learning jupyter notebook model deployment ai
Explore - Service reference
Bedrock
Build and scale generative AI applications. generative ai foundation models llm ai agent prompt engineering
Explore - Learning path
What the cloud actually changed
A company needs a server. In 2005 that meant a purchase order, a four-week lead time, a rack, and a machine sized for the busiest day of the year — which sat idle the other 364. Every constraint that shaped how software was built came from that fact.
Explore - Learning path
Regions and Availability Zones
Your application runs in one building. That building has one power feed, one cooling system, and one set of network cables. Everything AWS does about resilience is a response to the fact that buildings fail.
Explore - Learning path
Who is responsible when it breaks
A company put customer records in S3 and the records leaked. AWS was not breached. The bucket was configured to allow public reads. Knowing exactly where the line sits between AWS’s job and yours is the difference between a secure account and a news story.
Explore - Learning path
How the bill is actually calculated
The most common first AWS bill shock is not compute. It is a NAT Gateway nobody remembers creating, charging by the hour and again by the gigabyte, in three Availability Zones.
Explore - Learning path
Identity and permissions
Almost every AWS error a beginner hits is an IAM error wearing a different hat. AccessDenied on a bucket, a Lambda that cannot write logs, an EC2 instance that cannot reach a database — the same evaluation logic decides all three.
Explore - Learning path
Object storage
Marketing wants a landing page live today with no servers. Finance wants every resource attributable. Legal wants overwrites recoverable. All three are S3 settings, and two of them are off by default for good reasons.
Explore - Learning path
Virtual machines
An instance launches, reports "running", and refuses every connection. Nothing is broken. Four separate things — a subnet route, a security group, a network ACL, and a key pair — all have to agree before a packet reaches your process.
Explore - Learning path
The network your resources live in
Public and private subnets are not a setting. A subnet is "public" only because its route table has a path to an internet gateway. Once that clicks, NAT gateways, bastion hosts and VPC endpoints all stop being memorised facts.
Explore - Learning path
Load balancing and scaling
Traffic triples at 09:00 every weekday and the site falls over at 09:02. Adding a bigger instance fixes it until the day traffic quadruples. Horizontal scaling is a different shape of answer, and it forces your application to become stateless.
Explore - Learning path
Managed relational databases
The database is the part you least want to run yourself and the part hardest to move later. Multi-AZ and read replicas look similar in the console and solve completely different problems.
Explore - Learning path
Serverless compute
No servers to patch, and a bill of zero when nothing runs. In exchange you accept a 15-minute ceiling, cold starts, and the fact that your function may run twice for one event.
Explore - Learning path
HTTP in front of your code
A Lambda function is not reachable from the internet. Something has to terminate TLS, route paths, authenticate callers and throttle abuse before your code ever runs.
Explore - Learning path
Key-value at any scale
DynamoDB is fast and effectively unlimited, provided every query you will ever run is known before you design the table. Get the key schema wrong and the fix is a migration, not an index.
Explore - Learning path
Decoupling with queues and events
Service A calls service B directly. B goes down for ten minutes and A goes down with it. A queue between them turns an outage into a delay — and introduces a set of problems you now have to think about explicitly.
Explore - Learning path
Knowing what is happening
The site is slow. You have no metric that says so, no log you can search, and no alarm that fired. Observability is not a service you switch on afterwards; it is a design decision.
Explore - Learning path
Infrastructure as code
You built the environment by clicking. Now build it again, identically, in another Region, and prove nothing drifted. Clicking does not scale to that question.
Explore - Learning path
Consistency and partitioning
Two users read the same record a millisecond apart and get different answers. Nothing is broken. Understanding why is the difference between designing a distributed system and being surprised by one.
Explore - Learning path
Failure domains and blast radius
Every architecture decision is also a decision about what fails together. A single account, a single Region, a single security group, a shared database — each is a boundary, and boundaries are the only thing that limits damage.
Explore - Learning path
Retries, idempotency and backpressure
A downstream service slows down. Every caller retries. The retries multiply the load, which slows it further. Retry logic written without thought is the most reliable way to turn a blip into an outage.
Explore - Learning path
Caching, and what it really costs
A cache turns a latency problem into a correctness problem. Every layer you add buys speed and money and pays for it in staleness, and the hit rate decides whether the trade was worth making.
Explore - Learning path
Defence in depth
One control is a single point of failure for security exactly as much as one server is for availability. Real accounts layer identity, network, encryption and detection so that any one mistake is survivable.
Explore - Learning path
Designing to a budget
Two architectures both meet the requirements. One costs four times the other. Cost is a design constraint that shows up in every diagram — data transfer paths, instance lifecycles, storage classes — and almost never in the diagram itself.
Explore - Visual lesson
How AWS decides whether to allow a call
Every AccessDenied in AWS is the same algorithm returning a decision you did not expect. Once you can run that algorithm in your head, permission errors stop being mysterious.
Explore - Visual lesson
How renting virtual hardware changes software design
Moving to the cloud changes software from a procurement problem to an architecture decision, forcing engineers to design for dynamic capacity.
Explore - Visual lesson
How AWS physical locations affect latency and reliability
AWS partitions its global footprint into independent Regions and Availability Zones to let you trade speed, cost, and disaster recovery.
Explore - Visual lesson
Where AWS responsibilities end and your responsibilities begin
AWS protects the underlying infrastructure of the cloud, but securing your data and configuring your services remains entirely your job.
Explore - Visual lesson
How your monthly AWS bill is calculated
AWS charges for resources using four basic vectors — compute time, storage volume, request count, and data transfer — creating bills that accumulate even when servers are idle.
Explore - Visual lesson
Why moving data across AWS networks costs money
AWS data transfer charges depend heavily on network direction, boundary crossing, and routing, making network topology a major driver of overall cost.
Explore - Visual lesson
How AWS manages temporary access
An access key committed to a repository is a security breach waiting to happen, but learning to use temporary credentials through roles completely removes this risk.
Explore - Visual lesson
How S3 decides whether to allow access
Even a perfectly configured bucket policy will block access if account-level overrides are active, but understanding S3 evaluation logic makes debugging permission errors straightforward.
Explore - Visual lesson
How S3 achieves extreme durability
Replicating data across multiple data centres guarantees that it will survive hardware failures, but protecting against user error requires understanding S3 versioning and lifecycle mechanics.
Explore - Visual lesson
How storage classes trade cost for performance
Selecting the cheapest storage tier is a false economy if you ignore how often your code reads the data, but learning how retrieval fees work lets you optimise S3 costs without sacrificing performance.
Explore - Visual lesson
How EC2 instances manage execution state
Renting a virtual machine means managing its lifecycle, but knowing what survives a stop, start, or termination helps you design resilient compute while keeping costs down.
Explore - Visual lesson
How EBS provides persistent block storage
Block storage volumes are physically pinned to single data centres to ensure low latency, but understanding snapshots and IOPS limits lets you move and scale compute storage safely.
Explore - Visual lesson
How VPC routing creates public subnets
Public and private subnets are defined by route tables rather than simple checkboxes, but mastering subnet configuration and IP limits lets you design a secure, scalable network boundary.
Explore - Visual lesson
How Security Groups and NACLs defend networks
Securing a network requires firewalls at both the instance and subnet levels, but understanding statefulness and rule orders prevents mysterious connection timeouts.
Explore - Visual lesson
How NAT and endpoints connect private networks
Keeping database instances private keeps them secure, but routing their traffic through NAT gateways or free endpoints determines whether your design is cost-effective.
Explore - Visual lesson
How load balancers route incoming traffic
Clients need a single entry point that survives server failures. A load balancer decouples incoming requests from individual backend instances to keep systems online.
Explore - Visual lesson
Why scaling out takes longer than expected
Adding servers automatically is not instantaneous, and understanding the delay between a metric spike and a running application is key to preventing system outages.
Explore - Visual lesson
What a health check actually proves
A health check determines whether a server is capable of processing work, but poorly designed checks can easily trigger self-inflicted outages.
Explore - Visual lesson
How RDS keeps databases online
High availability is not performance, and understanding how RDS replicates data and manages DNS failover is critical for designing resilient database tiers.
Explore - Visual lesson
Why read replicas are not failover targets
Replicating data asynchronously allows you to scale read workloads, but read replicas cannot protect your application from a database write failure.
Explore - Visual lesson
How Lambda executes your code
Serverless compute removes the operating system layer, but forces you to work within memory and execution time ceilings.
Explore - Visual lesson
Why concurrency is the true unit of Lambda capacity
Lambda scales by running multiple instances of your function in parallel, and managing this concurrency prevents database overload and application throttling.
Explore - Visual lesson
How API Gateway protects your backend
An API Gateway is the entry point that manages HTTP routing, security, and throttling before your serverless code ever executes.
Explore - Visual lesson
How DynamoDB keys locate your data
Your database key schema is your primary query interface, and choosing the wrong partition key leads to hot partitions and serverless scalability limits.
Explore - Visual lesson
Why a DynamoDB scan will destroy your budget
Scans read entire tables and consume massive capacity, whereas indexes provide efficient and cost-effective access paths to your data.
Explore - Visual lesson
How queues buffer unpredictable spikes
Placing a queue between services converts immediate outages into manageable processing delays and protects downstream systems.
Explore - Visual lesson
How fanout pattern distributes messages
Publishing an event to a topic allows multiple downstream queues to receive the message independently, isolating failures and scaling systems.
Explore - Visual lesson
How EventBridge routes events by pattern
EventBridge decouples systems by routing JSON payloads based on their content, but unmatched events vanish silently unless you configure archives or dead-letter queues.
Explore - Visual lesson
Why you should alarm on symptoms rather than causes
Alarming on every CPU spike leads to paging fatigue, whereas targeting end-user symptoms ensures you only wake engineers for real outages.
Explore - Visual lesson
How CloudFormation manages infrastructure as state
Describing infrastructure in code ensures consistency, but executing updates requires evaluating changesets to prevent accidental resource deletion.
Explore - Visual lesson
How databases trade consistency for speed
Distributed systems must choose between immediate data consistency and low read latency, forcing application design to handle stale reads.
Explore - Visual lesson
How failure domains limit damage
Partitioning resources into isolated logical and physical boundaries prevents localised component failures from causing global system outages.
Explore - Visual lesson
How retries multiply system load
Uncoordinated client retries amplify downstream database and service failures, transforming minor network glitches into permanent system outages.
Explore - Visual lesson
How idempotency handles duplicates
Since network failures make duplicate messages inevitable, applications must use idempotency keys and conditional writes to ensure safe execution.
Explore - Visual lesson
How caching hit rates dictate origin load
A cache trades data consistency for read throughput, but a slight drop in the hit rate can double the traffic hitting your database.
Explore - Visual lesson
How security layers defend in depth
Layering network, identity, and cryptographic controls ensures that a single misconfiguration does not lead to a catastrophic security breach.
Explore - Visual lesson
How cost functions as a design constraint
Cost is an architectural constraint that must be engineered into your diagram, as data transfer and idle resources dictate the monthly bill.
Explore - Problem guide
EC2 SSH connection timed out: trace the broken hop
A timeout means the connection did not complete. Check the instance, public address, subnet route, security group and both NACL directions before changing credentials.
Explore - Problem guide
Private subnet has no internet: check the NAT path
A private instance needs a usable outbound path. For public NAT, verify the private route to NAT, NAT state and VPC, its public subnet route to an attached internet gateway, and both subnet ACLs.
Explore - Problem guide
S3 AccessDenied: identify the permission boundary
A 403 is an access decision, not proof that the object is missing. Identify the caller, action and resource; then check matching allows, explicit denies and the applicable public-access or encryption controls.
Explore - Problem guide
IAM explicit deny: why another allow cannot fix it
A matching explicit deny overrides an allow. First identify which statement applies to the exact action, resource and principal; adding a broader allow cannot cancel that deny.
Explore - Problem guide
SQS backlog keeps growing: compare arrival and drain rates
A queue grows when incoming work exceeds completed work over time. Separate arrival rate, worker capacity and retries before increasing concurrency or changing visibility timeout.
Explore - Problem guide
Retry storms: when recovery logic multiplies the outage
Retries consume capacity too. Bound attempts, use backoff and jitter, and make side effects idempotent so failure recovery does not amplify load or duplicate work.
Explore - Comparison
Security groups vs NACLs: follow the request and reply
Security groups apply to associated resources and are stateful. NACLs filter subnet-boundary traffic, support allow and deny rules, and require separate permission for return traffic.
Explore - Comparison
NAT gateway vs internet gateway: two different paths
An internet gateway provides a VPC internet path for appropriately addressed resources. Public NAT lets private IPv4 resources initiate internet connections through a translated public address.
Explore - Comparison
IAM users vs roles: choose credentials for the workload
IAM users identify account identities that can have long-lived credentials. Roles are assumed to obtain temporary credentials. Prefer suitable federation for people and roles for workloads.
Explore - Comparison
SQS vs SNS vs EventBridge: queue, publish or route?
Use SQS to buffer work for consumers, SNS to publish to subscribers, and EventBridge to route events using rules. They can be combined when a system needs both routing and durable buffering.
Explore - Experiment
Network path debugger
Break and repair IPv4 routes, NAT paths and return traffic.
Explore - Experiment
IAM permission evaluator
See allows, explicit denies and boundary restrictions for one request.
Explore - Experiment
S3 anonymous access experiment
Test policy scope and public-read restrictions.
Explore - Experiment
Capacity and scaling model
Compare demand with per-instance capacity and a scaling ceiling.
Explore - Experiment
Failure domain experiment
Take one zone offline and inspect remaining capacity.
Explore - Experiment
Queue backlog model
Step incoming work and worker drain capacity over time.
Explore - Experiment
Replica lag experiment
Write a value and compare primary and delayed replica reads.
Explore - Experiment
Retry amplification calculator
Count worst-case downstream attempts across retrying layers.
Explore - Experiment
Cache economics calculator
Change hit rate, origin cost and fixed cache cost.
Explore - CLF-C02
AWS Cloud Practitioner preparation
Concept preparation. Implementation, architecture design and troubleshooting are outside the stated target tasks. This collection covers selected topics, not the full exam.
Explore - SAA-C03
AWS Solutions Architect Associate preparation
Use constraints to choose an architecture, then test the relevant concept. Selected explanations and local experiments are available; this is not full exam coverage.
Explore - SOA-C03
AWS CloudOps Engineer Associate preparation
An operational topic map for selected local practice. Domain labels here are learning groups, not a complete weighted exam blueprint. Verify the official current guide before planning your exam.
Explore