AWS SAA-C03 Trilingual Cheat Sheet Digest

2026-07-30

This is a trilingual (Chinese / Japanese / English) study digest for the AWS SAA-C03 exam, covering all ten SAA-C03 domains. Each section carries πŸ”‘ key numbers, ⚠️ exam traps, and πŸ“Š decision trees for quick review during final prep. The must-memorize high-frequency numbers include: Lambda's 15-minute max runtime, 10GB memory ceiling, 10GB /tmp space, and 6MB synchronous payload limit; DynamoDB's 400KB max item size and 3,000 RCU / 1,000 WCU per-partition limits; S3's eleven nines of durability and Glacier Deep Archive's 180-day minimum storage duration; EBS gp3's baseline of 3,000 IOPS plus 125 MB/s; and RDS's 1–35 day backup retention window. One major change worth flagging: the Snow Family has been largely retired β€” Snowmobile (the 100PB truck) was retired in March 2024, Snowcone and the legacy 80TB Snowball Edge were discontinued on November 12, 2024, and the entire Snow Family stopped accepting new customers as of November 7, 2025. Exam question banks may still use the historical '80TB/100PB' figures, and the traditional answers still apply on test day, but in real-world work DataSync should be the preferred choice.

1. Compute

EC2 pricing models: On-Demand has no commitment and is the most expensive, suited for short-term or unpredictable workloads. Reserved Instances require a 1- or 3-year commitment β€” Standard RI offers up to roughly 72% discount but locks the instance family, while Convertible RI offers roughly 54–66% but lets you change instance family, OS, or term. Spot Instances offer up to roughly 90% off but can be interrupted with only 2 minutes' notice, making them suitable for stateless, fault-tolerant, or batch workloads. Savings Plans commit to a dollar amount of usage per hour β€” Compute Savings Plans are the most flexible, covering EC2/Fargate/Lambda, while EC2 Instance Savings Plans offer a higher discount but lock the instance family. Separately, a Dedicated Host exposes the physical server's sockets/cores, meeting bring-your-own-license (BYOL) and compliance needs billed per physical core; a Dedicated Instance dedicates the hardware to your account but doesn't expose physical cores, so it can't support per-core BYOL.

Key numbers & exam traps: Spot interruption notice is 2 minutes; RI/SP terms are 1 or 3 years; Standard RI discounts up to ~72%, Convertible ~54–66%, Spot up to ~90%. Common trap: when a question mentions needing to see physical sockets/cores for BYOL, pick Dedicated Host, not Dedicated Instance. For 'fault-tolerant batch work, lowest cost' pick Spot; for 'flexible savings across EC2/Fargate/Lambda' pick Compute Savings Plans (not Standard RI, which locks the instance family). Also note Standard RI can change AZ/instance size (within the same family)/scope, but cannot change instance family β€” use Convertible RI for that. Decision flow: unpredictable short-term β†’ On-Demand; stable predictable 24/7 β†’ Standard RI/SP; interruptible fault-tolerant β†’ Spot; need visible physical cores or BYOL compliance β†’ Dedicated Host.

Lambda: serverless, event-driven function execution. Max runtime is 15 minutes (900 seconds); memory ranges 128MB–10,240MB (CPU scales linearly with memory); /tmp storage is 512MB–10GB; synchronous invocation payload limit is 6MB, asynchronous is 256KB; default account concurrency is 1000 (a soft limit that can be raised). Cold start is the latency from initializing the execution environment on first invocation or scale-out. Provisioned Concurrency pre-warms environments to eliminate cold starts (costs extra); Reserved Concurrency reserves/caps a concurrency quota for a specific function (prevents one function from consuming the whole account's concurrency, itself free). Layers share dependencies, up to 5 per function. Key numbers: 15 min | 10GB memory | 10GB /tmp | 6MB sync / 256KB async payload | default 1000 concurrency | up to 5 layers | deployment package limits of 50MB (zip upload) / 250MB (unzipped) / 10GB (container image). Exam traps: tasks over 15 minutes need Step Functions/ECS/Batch instead β€” Lambda's own timeout can't be extended. 'Eliminate cold-start latency' β†’ Provisioned Concurrency; 'guarantee/limit concurrency for a specific function' β†’ Reserved Concurrency β€” don't confuse the two. Attaching Lambda to a VPC increases cold starts due to ENI creation; for many DB connections, use RDS Proxy for connection pooling.

ECS: Fargate vs EC2 launch type: ECS is container orchestration. Fargate is serverless β€” AWS manages the underlying infrastructure, billed per vCPU/memory, no ops overhead, good for bursty workloads or when you don't want to manage servers. The EC2 launch type requires you to manage the EC2 cluster yourself, is cheaper at large stable scale, supports GPUs/special instance types, and gives you control over the host. EKS is managed Kubernetes; ECR is the container image registry. Exam trap: 'don't want to manage servers / bursty workload' β†’ Fargate; 'need GPUs, large stable workload to save cost, need host-level access' β†’ EC2 launch type.

Auto Scaling: ASG automatically scales EC2 instance count. Lifecycle Hooks pause an instance during launch/terminate so custom actions (e.g., installing software, exporting logs) can run. Scaling policies include Target Tracking (recommended, e.g., maintain 50% average CPU), Step Scaling, Simple Scaling (with a cooldown), and Scheduled scaling. Cooldown pauses further Simple Scaling actions for a period (default 300 seconds) to prevent flapping. Exam trap: 'export logs / graceful shutdown before instance termination' β†’ Lifecycle Hook (fires in the terminating state). Health checks default to EC2 status checks only; for application-level health, enable ELB health checks.

Elastic Beanstalk: a PaaS β€” upload your code and it auto-deploys EC2+ASG+ELB+monitoring, while you still own and can control the underlying resources. Good for quickly deploying web apps when developers don't want to touch infrastructure directly. Deployment strategies: All at once, Rolling, Rolling with additional batch, Immutable, and Blue-Green.

2. Storage

S3 storage classes: S3 Standard has no minimum storage duration, millisecond retrieval, suited for frequently accessed data, spread across 3+ AZs. Standard-IA has a 30-day minimum with millisecond retrieval, for infrequently accessed data that still needs fast retrieval. One Zone-IA also has a 30-day minimum but lives in a single AZ, for infrequent, re-creatable data. Glacier Instant Retrieval has a 90-day minimum with millisecond retrieval, for data accessed roughly quarterly but needing instant access. Glacier Flexible Retrieval has a 90-day minimum with retrieval taking minutes to 12 hours, for backup/DR data accessed roughly twice a year. Glacier Deep Archive has a 180-day minimum with 12–48 hour retrieval, for annual-access compliance archives. Intelligent-Tiering has no minimum duration and millisecond retrieval, for unknown access patterns that should auto-tier. All classes except One Zone-IA span 3+ AZs.

Glacier Flexible Retrieval has three speed tiers: Expedited (1–5 min), Standard (3–5 hours), Bulk (5–12 hours, free). Deep Archive has no Expedited β€” only Standard (12 hours) and Bulk (48 hours). All classes carry eleven nines (99.999999999%) durability. Other key S3 features: Lifecycle Policy auto-transitions or expires objects; Replication splits into cross-region CRR and same-region SRR, both requiring versioning; Versioning keeps multiple object versions to prevent accidental deletion/overwrite; MFA Delete requires MFA to delete versions or disable versioning, enabled only by the root user via CLI; Presigned URLs grant temporary access to private objects; Multipart Upload splits large files β€” recommended above 100MB, required above 5GB, with a 5TB max object size. Key numbers: max object size 5TB; single PUT limit 5GB (multipart required beyond that); multipart recommended threshold 100MB; One Zone-IA and Standard-IA share a 30-day minimum; Glacier Instant Retrieval's minimum object is 128KB; durability is eleven nines across the board; Standard availability is 99.99%. Exam traps: 're-creatable, infrequent, cost-sensitive, tolerant of single-AZ loss' β†’ One Zone-IA; 'archived but occasionally needs millisecond reads (e.g., medical imaging)' β†’ Glacier Instant Retrieval (not Standard-IA or Flexible); 'lowest cost, can wait 12+ hours' β†’ Deep Archive; 'unpredictable access pattern, don't want manual management' β†’ Intelligent-Tiering (no retrieval fee, no minimum duration, but a small monitoring fee). Deleting IA/Glacier objects early still incurs the minimum-duration charge.

EBS volume types: gp3 (SSD) β€” up to 16,000 IOPS, up to 1,000MB/s throughput, baseline 3,000 IOPS + 125MB/s (independent of size, IOPS and throughput configurable separately), ~20% cheaper per GB than gp2 (~$0.08 vs $0.10/GiB-month, released Dec 2020) β€” the recommended general-purpose choice. gp2 (SSD) β€” up to 16,000 IOPS but only 250MB/s throughput, the legacy general type, performance relies on burst credits (3 IOPS/GB). io1 (SSD) β€” up to 64,000 IOPS, 1,000MB/s throughput, high performance, supports Multi-Attach. io2 Block Express (SSD) β€” up to 256,000 IOPS, 4,000MB/s throughput, 99.999% durability, for extreme-performance workloads like SAP HANA. st1 (HDD) β€” up to 500 IOPS, 500MB/s throughput, for large sequential I/O like logs/big data. sc1 (HDD) β€” up to 250 IOPS, 250MB/s throughput, the cheapest cold-data option. Snapshots are stored incrementally in S3 and can be copied cross-region/cross-account; encryption uses KMS across snapshots, volumes, and replicas, and volumes restored from encrypted snapshots are automatically encrypted.

Exam traps: small gp2 volumes rely on burst credits and drop to baseline (3 IOPS/GB) once exhausted β€” gp3 is the go-to exam answer. Multi-Attach is io1/io2 only and needs a cluster file system. EBS volumes can only attach to EC2 instances in the same AZ; moving across AZs requires snapshot-then-restore. Encrypting an unencrypted volume: snapshot β†’ copy with encryption enabled β†’ restore from the encrypted snapshot. EFS: managed NFS, POSIX-compliant, mountable by multiple EC2 instances across AZs simultaneously, elastically scaling. Storage classes: Standard (multi-AZ) vs One Zone (single AZ, cheaper). Performance modes: General Purpose (default, low latency) vs Max I/O (higher concurrency/throughput, slightly higher latency). Throughput modes: Bursting, Provisioned, Elastic; lifecycle management can auto-move files to IA to save cost. Exam trap: EFS is a Linux shared filesystem β€” for Windows shares use FSx for Windows, for HPC use FSx for Lustre. EBS is single-instance block storage and, aside from Multi-Attach, can't be shared across instances.

Snow Family (⚠️ major changes): as historical knowledge (still the traditional exam answer), Snowcone is ~8TB HDD / 14TB SSD, the most portable; the storage-optimized Snowball Edge is 80TB; Snowmobile is a truck holding up to 100PB. The rule for choosing a physical device is: when data volume is too large for network transfer to be practical, switch to physical devices (see bandwidth math in section 8). Real-world changes (2024–2025): Snowmobile was retired in March 2024; Snowcone and the legacy 80TB Snowball Edge were discontinued November 12, 2024; only the latest-generation 210TB Snowball Edge remains, and even it stopped accepting new customers as of November 7, 2025 β€” AWS now recommends DataSync/Data Transfer Terminal/Outposts instead. Exam trap: the SAA-C03 question bank may still present '80TB/100PB' options, and the traditional mapping still applies there β€” but be aware these devices are retired in reality, and AWS's current push is toward online transfer via DataSync.

Storage Gateway types: a hybrid cloud storage gateway giving on-premises access to cloud storage. File Gateway offers NFS/SMB interfaces, storing files in S3. Volume Gateway offers iSCSI block storage, split into Cached (primary data in S3, only hot data kept locally) and Stored (primary data stays on-prem, asynchronously backed up to S3). Tape Gateway (VTL) replaces physical tape backup infrastructure, storing data in Glacier. Exam trap: 'replace physical tape backup infrastructure' β†’ Tape Gateway; 'on-prem app needs low-latency access and primary data must stay local' β†’ Stored Volume.

3. Database

RDS: Multi-AZ provides HA/DR β€” data is synchronously replicated to a standby in another AZ, with automatic failover on primary failure (DNS endpoint unchanged), but the standby cannot be read. Read Replicas provide read scaling via asynchronous replication, are themselves readable, can span AZs/regions, support up to 5 replicas, and can be manually promoted to standalone databases. Automated backups take a daily full backup plus transaction logs every 5 minutes, with retention of 1–35 days (default 7), supporting PITR to any second. Manual snapshots are user-triggered and kept permanently (surviving even DB deletion). RDS Proxy provides connection pooling, reducing DB connection overhead β€” especially useful with Lambda β€” and speeding up failover. Parameter Groups manage DB engine configuration. Key numbers: backup retention 1–35 days (default 7, 0 disables it); transaction logs every 5 minutes; up to 5 read replicas (MySQL supports secondary replicas). Exam traps: Multi-AZ β‰  read scaling β€” the standby isn't readable, so use Read Replica for that. Read Replica β‰  automatic failover (requires manual promotion) β€” use Multi-AZ for HA; the two can be combined. Encrypting an existing unencrypted RDS instance requires snapshot β†’ encrypted copy β†’ restore (no in-place encryption). Needing OS/SSH access isn't possible with standard RDS β€” use RDS Custom or a self-managed EC2 database instead.

Aurora: AWS's proprietary database engine, MySQL/PostgreSQL-compatible, with the storage layer automatically maintaining 6 copies across 3 AZs. Global Database supports cross-region deployment with replica lag typically under 1 second (RPO ~1s, RTO ~1min typically), used for globally low-latency reads plus regional DR, with the write endpoint automatically pointing to the new primary after failover. Serverless v2 auto-scales on demand, measured in ACUs (roughly 2GiB memory each), ranging 0.5–128 ACU in 0.5 increments. Cluster endpoints: Writer (points to the primary), Reader (load-balances across replicas), and Custom (a custom subset), with up to 15 replicas. Exam trap: 'global users need low-latency reads plus regional DR' β†’ Aurora Global Database; 'unpredictable/intermittent load, want zero management' β†’ Serverless v2. Use the Writer endpoint for writes and the Reader endpoint for reads (don't send reads to Writer).

DynamoDB: a NoSQL key-value/document database, fully managed and serverless, with millisecond response times. Capacity modes: On-Demand (pay per request, good for unpredictable traffic) vs Provisioned (preset RCU/WCU, good for predictable traffic, can pair with Auto Scaling). Max item size is 400KB (larger items should go in S3 with only a pointer in DynamoDB). Per-partition limit is 3,000 RCU/1,000 WCU β€” partition keys need high cardinality and even distribution to avoid 'hot partitions'. DAX is an in-memory cache offering microsecond reads, purpose-built for DynamoDB and directly pluggable in. Global Tables provide multi-region multi-active replication. Streams provide change capture that can trigger Lambda. 1 RCU = one strongly consistent read (≀4KB) or two eventually consistent reads; 1 WCU = one write (≀1KB). Key numbers: 400KB max item | 3,000 RCU/1,000 WCU per partition | 1 RCU = 4KB strong read | 1 WCU = 1KB write | Query returns ≀1MB per call | BatchGetItem ≀16MB | default table quota 40,000 RCU/40,000 WCU. Exam traps: 'need microsecond read caching' β†’ DAX (not ElastiCache); 'multi-region low-latency reads/writes' β†’ Global Tables; 'storing objects over 400KB' β†’ store in S3, keep only a reference/pointer in DynamoDB; hot-partition issues β†’ optimize the partition key (high cardinality, even distribution), avoid low-cardinality keys; 'respond to DynamoDB data changes and trigger processing' β†’ Streams + Lambda.

ElastiCache: Redis vs Memcached: Redis (including Valkey) supports rich data structures (strings/lists/sets/sorted sets/hashes), persistence (RDB/AOF), replication/Multi-AZ/failover, runs on a single main thread, and supports Pub/Sub β€” good for leaderboards, session storage, durable caching, messaging. Memcached only supports simple string key-values, has no persistence (data lost on restart), no replication/Multi-AZ/failover, runs multi-threaded, and has no Pub/Sub β€” good for simple caching and maximum-throughput horizontal scaling. Exam trap: 'real-time leaderboard (sorted sets)/geospatial data/need persistence and failover' β†’ Redis; 'simplest cache, multi-threaded max throughput, data loss acceptable' β†’ Memcached; 'microsecond caching for DynamoDB' β†’ DAX (not ElastiCache).

Other databases at a glance: Redshift is a petabyte-scale data warehouse with columnar storage, for OLAP analytics/complex queries/BI. Neptune is a graph database, for social networks/recommendations/knowledge graphs/fraud detection. DocumentDB is a MongoDB-compatible document database. Athena runs serverless SQL directly on S3 data, paired with QuickSight for BI visualization. Exam trap: 'run SQL directly on S3 data, no preloading' β†’ Athena; 'data warehouse analytics on historical data' β†’ Redshift; 'highly relational queries (friend-of-a-friend)' β†’ Neptune.

4. Networking

VPC basics: A VPC is a virtual private cloud. Subnets (public/private) each bind to a single AZ; Route Tables handle traffic forwarding; an IGW gives public subnets internet access. A NAT Gateway (managed, HA) lets private subnets reach the internet outbound without accepting inbound connections, managed by AWS and billed hourly plus data processed. A NAT Instance (self-run EC2) requires self-management and self-built HA, can double as a bastion host, is cheaper but a legacy approach. Exam trap: NAT Gateway must sit in a public subnet, with private subnet routes pointing to it; one NAT Gateway per AZ is needed for HA; NAT Instances need source/destination check disabled.

Security Group vs NACL: Security Groups are stateful β€” return traffic is automatically allowed β€” operate at the instance/ENI level, support only Allow rules, can reference other security groups, are evaluated as a whole, and default to denying all inbound / allowing all outbound. NACLs are stateless β€” return traffic needs an explicit rule β€” operate at the subnet level, support both Allow and Deny, reference only CIDR blocks, are evaluated in rule-number order, and the default NACL allows all traffic. Exam trap: 'block a specific malicious IP' can only be done with NACL Deny rules (security groups have no deny). Since NACLs are stateless, opening an inbound rule also requires opening the outbound ephemeral port range (1024–65535) or return traffic gets blocked. Security groups can reference another security group (e.g., 'allow only traffic from an ALB's security group') β€” NACLs cannot. The default security group denies inbound while the default NACL allows everything β€” but a custom NACL defaults to denying everything.

VPC connectivity: Peering vs Transit Gateway vs PrivateLink: VPC Peering directly connects two VPCs, is non-transitive (A-B and B-C don't imply A-C), disallows overlapping CIDRs, has no processing fee within the same region (only cross-AZ transfer), suited for a few VPCs. Transit Gateway is a hub-and-spoke connecting many VPCs/on-premises (VPN/DXGW), supports transitive routing, billed per attachment (~$0.05/hr per attachment + $0.02/GB processed, e.g. us-east-1), suited for 3+ VPCs. PrivateLink (VPC Endpoint Service) exposes only a single service rather than the whole network, accessed via the consumer's ENI, supports overlapping CIDRs, follows a zero-trust model, suited for SaaS/service-sharing. Exam trap: with overlapping CIDRs, neither Peering nor Transit Gateway works β€” only PrivateLink. VPC Peering is non-transitive β€” fully connecting three VPCs needs three peering connections or a Transit Gateway. For high-bandwidth pairs of VPCs, Peering can save on Transit Gateway processing fees.

Direct Connect vs VPN: Direct Connect (DX) uses a dedicated physical line β€” high, consistent bandwidth (1/10/100Gbps), low consistent latency β€” but takes weeks to months to set up and costs more. Site-to-Site VPN connects via IPsec tunnels over the internet β€” bandwidth limited by internet conditions, latency fluctuates β€” but sets up in minutes to hours and costs less. Exam trap: 'need a hybrid connection immediately' β†’ VPN (DX takes weeks); 'need stable high bandwidth and low latency' β†’ DX. For DX backup/encryption, combine DX+VPN β€” DX itself isn't encrypted, so use VPN over DX when encryption is required.

Route 53 routing policies: Simple (single resource); Weighted (proportional traffic split, for A/B testing or canary releases); Latency (routes to the lowest-latency region); Failover (active-passive with health checks); Geolocation (routes by user location, for compliance/localization); Geoproximity (routes by resource location with a bias shift, needs Traffic Flow); Multivalue Answer (returns up to 8 healthy records at random, simple load balancing). Exam trap: Geolocation is based on where the user is (compliance/language), Geoproximity is based on where the resource is with bias adjusting traffic share (needs Traffic Flow) β€” don't confuse them. 'Lowest latency' β†’ Latency (not Geoproximity). Multivalue isn't an ELB replacement, but returning multiple healthy IPs gives simple load balancing.

CloudFront: a CDN with edge caching. Origins can be S3/ALB/an HTTP server; Behaviors route by path or set cache policies. Signed URLs authorize access to a single file (e.g., installer downloads), also useful when the client doesn't support cookies. Signed Cookies authorize access to multiple files (e.g., an entire subscription area or HLS video streams), useful when you don't want to change URLs. OAC/OAI restrict an S3 origin to only be accessible via CloudFront. Exam trap: multi-file or video-stream authorization β†’ Signed Cookie; single file β†’ Signed URL; if both apply, Signed URL takes precedence. Restricting direct S3 access β†’ OAC (new) / OAI (legacy).

Load balancer types (ELB): ALB operates at Layer 7 (application), supports HTTP/HTTPS, path/host-based routing, WebSocket, suited for containers and WAF integration. NLB operates at Layer 4 (transport), supports TCP/UDP/TLS, offers ultra-high performance and ultra-low latency, static IPs, and can handle millions of connections. CLB operates at Layers 4 and 7, a legacy option not recommended for new projects. GWLB operates at Layer 3 (network), used to deploy third-party virtual security appliances. Exam trap: 'static IP/extreme performance/TCP or UDP gaming' β†’ NLB; 'URL path or hostname-based routing, microservices, containers' β†’ ALB. Needing to preserve client source IP at the backend β€” NLB preserves it by default.

Global Accelerator vs CloudFront: CloudFront is a CDN caching content (HTTP/HTTPS) at the edge, for static/dynamic web content and video streaming. Global Accelerator is a network-layer (L4) acceleration service offering 2 static anycast IPs, routed over the AWS backbone to the nearest healthy endpoint, supporting TCP/UDP with fast cross-region failover, suited for games, IoT, VoIP, and other non-HTTP or static-IP-requiring use cases. Exam trap: 'static IP, non-HTTP (gaming/VoIP), fast multi-region failover' β†’ Global Accelerator; 'cache websites/video, WAF integration' β†’ CloudFront. Both leverage the AWS backbone network and Shield protection.

5. Security & IAM

IAM basics: Users (long-term credentials), Groups (bulk permission grants), Roles (temporary credentials, assumed by services/cross-account/federated identities), Policies (JSON permission definitions). Inline policies are embedded 1:1 in an identity and deleted with it; managed policies are reusable, either AWS-managed or customer-managed. The guiding principle is least privilege, and an explicit Deny always wins over any Allow.

Trust Policy vs Permission Policy: A Trust Policy defines 'who can assume this role' (the Principal), and only attaches to a Role. A Permission Policy defines 'what this role/identity can do.' A role needs both. STS: AssumeRole obtains temporary credentials (default 1-hour validity), used for cross-account access, EC2/Lambda roles, and federation. Temporary credentials consist of an AccessKey, SecretKey, and SessionToken.

KMS: Envelope Encryption works by having KMS not directly encrypt bulk data β€” instead, a data key (DEK) is generated to encrypt the data, and that DEK is itself encrypted ('wrapped') by a CMK/KMS key (KEK), with the encrypted DEK stored alongside the ciphertext. To decrypt, the encrypted DEK is sent to KMS to be exchanged back for the plaintext DEK. The CMK itself never leaves KMS. AWS-managed keys (e.g. aws/s3) can't have their key policy modified or be manually rotated β€” AWS rotates them automatically each year. Customer-managed keys (CMK) let you control the key policy, enable/disable the key, schedule deletion (7–30 day waiting period), and enable automatic rotation. Auto-rotation is on a 1-year cycle (symmetric keys only β€” asymmetric/HMAC/imported key material aren't supported). AWS changed AWS-managed key rotation to roughly annual in May 2022, and added on-demand rotation support starting May 2024; rotation doesn't re-encrypt existing data β€” KMS retains old key material to decrypt older data. A Key Policy is a resource-based policy that every key must have exactly one of, and is the primary access control mechanism; IAM policies alone aren't sufficient for authorization (unless the key policy explicitly enables IAM), and cross-account access requires bidirectional authorization (both key policy and IAM policy). S3 encryption options: SSE-S3 (fully managed by S3, AES-256), SSE-KMS (managed by KMS, auditable, cross-account capable β€” watch KMS request quotas and enable S3 Bucket Keys to reduce call volume), SSE-C (customer supplies the key on every request), and client-side encryption (customer encrypts before upload).

KMS exam traps (classic topics): copying an encrypted snapshot cross-region requires re-encryption β€” KMS keys are regional, the source key can't be used in another region, and the copy operation must specify the destination region's key (--kms-key-id). A snapshot encrypted with the default AWS-managed key can only be shared within the same account; sharing cross-account or cross-region requires first copying it as a customer-managed CMK-encrypted version and updating the key policy for authorization. 'Need to audit who used the key, share cross-account, or use a custom rotation schedule' β†’ customer-managed CMK + SSE-KMS. SSE-KMS under high-concurrency GET/PUT can trigger KMS throttling (ThrottlingException) β€” enable S3 Bucket Keys. 'Customer wants full key control, AWS shouldn't store the key' β†’ SSE-C or client-side encryption. Encryption: in-transit vs at-rest: in-transit encryption uses TLS/SSL for network transport (HTTPS, ACM certificates); at-rest encryption is storage-layer encryption (KMS, EBS/S3/RDS encryption).

Security monitoring services compared: CloudTrail answers 'who did what' β€” records all API calls, an API audit service. Config answers 'what changed' β€” records resource configuration change history and pairs with compliance rules. GuardDuty answers 'is there a threat' β€” analyzes CloudTrail/VPC Flow/DNS logs for intelligent threat detection. Inspector answers 'is there a vulnerability' β€” scans EC2/Lambda/containers for software vulnerabilities. Macie answers 'is there sensitive data in S3' β€” uses ML to detect PII/sensitive data in S3. Exam trap: these five are a favorite exam topic, so distinguish them clearly β€” 'detect a compromised instance/anomalous API calls/malicious IPs' β†’ GuardDuty; 'check resources against a config baseline/audit config changes' β†’ Config; 'scan EC2 for vulnerabilities/CVEs' β†’ Inspector; 'find credit card numbers or PII in S3' β†’ Macie; 'find out who deleted a resource' β†’ CloudTrail. Security Hub aggregates findings from these services (except Config itself).

WAF/Shield: WAF is a Layer-7 web application firewall against SQL injection/XSS, based on rules/IP/geography/rate limiting, attachable to ALB/CloudFront/API Gateway. Shield Standard is free, providing automatic L3/4 DDoS protection for all customers. Shield Advanced is paid (~$3,000/month, 1-year commitment), offering enhanced DDoS protection, cost-spike protection, DDoS Response Team access (SRT, requires Business/Enterprise Support), and includes WAF. Exam trap: 'defend against SQL injection, URL-based or rate limiting' β†’ WAF; 'large-scale DDoS, need expert support and cost protection' β†’ Shield Advanced.

Cognito: User Pools handle authentication (who you are) β€” a managed user directory handling sign-up/sign-in/MFA, ultimately returning JWT tokens (ID/Access/Refresh). Identity Pools (federated identities) handle authorization (which AWS resources you can access) β€” they don't store users, but exchange tokens (from User Pools or third parties like Google/Facebook) via STS for temporary AWS IAM credentials to access S3, DynamoDB, etc. Exam trap: 'sign-in, user directory, issuing JWTs' β†’ User Pool; 'exchange tokens for temporary AWS credentials to access S3/DynamoDB' β†’ Identity Pool. Full flow: User Pool authentication β†’ JWT β†’ Identity Pool exchanges via STS β†’ temporary credentials β†’ access AWS resources.

SCP vs Permission Boundary: SCPs operate at the Organizations level, setting a maximum/guardrail on member account permissions (they don't grant permissions, only restrict), affecting the entire account (root included). Permission Boundaries set a permission ceiling for an individual IAM user or role β€” the effective permission is the intersection of the identity policy and the boundary. Exam trap: SCPs never actively grant permissions, only cap them; even if an IAM policy grants Allow, the action is still denied if the SCP doesn't permit it. Effective permission = SCP ∩ identity policy ∩ permission boundary (intersection), and an explicit Deny always wins.

6. Application Integration

SQS: a managed message queue for decoupling, using a polling model. Standard queues offer near-unlimited throughput, at-least-once delivery, but only best-effort ordering (possible duplicates/reordering β€” consumers need to be idempotent). FIFO queues guarantee strict ordering and exactly-once processing (a 5-minute dedup window), with throughput of 300 msg/s (no batching) or 3,000 msg/s (batched), higher in high-throughput mode. Visibility Timeout is how long a retrieved message stays hidden (default 30s, max 12 hours) β€” if not deleted before processing finishes, it reappears. A DLQ catches messages that fail beyond maxReceiveCount (must match the source queue's type). Long polling (WaitTimeSeconds > 0) reduces empty responses and API cost (max 20s). Key numbers: max message size 256KB (larger needs the Extended Client + S3); max retention 14 days (default 4); visibility timeout default 30s/max 12h; FIFO throughput 300/3,000 (batched) msg/s; long polling max 20s; delay queues max 15 min; in-flight messages ~120,000 (both Standard and FIFO β€” AWS raised the FIFO in-flight limit from 20,000 to ~120,000 in November 2024).

Exam trap: 'strict ordering, no duplicate processing allowed (orders, transactions)' β†’ FIFO; 'maximum throughput, duplicates tolerable' β†’ Standard (consumers must be idempotent). If Lambda processing takes 45s but visibility timeout is set to only 30s, the message reappears and gets processed twice! Set the visibility timeout to several times the processing time (6x is a good rule of thumb). DLQ type must match the source queue (FIFO queue needs a FIFO DLQ). To reduce empty-poll costs, use long polling. SNS: a publish/subscribe messaging service using a push model. Fan-out means one SNS topic delivers simultaneously to multiple SQS queues, Lambda functions, or HTTP subscribers. Core concepts are Topics and Subscriptions (email/SMS/SQS/Lambda/HTTP), often combined with SQS (SNS+SQS fan-out). Exam trap: 'one message needs to be distributed to multiple systems/queues at once' β†’ SNS fan-out to multiple SQS queues (rather than each system polling independently); SNS+SQS combined achieves both buffering and durability.

EventBridge: an event bus with more powerful routing than SNS. Supports scheduled events (cron/rate), matches event patterns via Rules, and routes to Targets (Lambda/SQS/SNS, etc.). Supports SaaS/third-party event sources and a Schema Registry. Exam trap: 'route based on event content, connect SaaS events, replace CloudWatch Events' β†’ EventBridge; 'simple scheduled Lambda trigger' can use EventBridge Scheduler (formerly CloudWatch Events). Step Functions: a serverless workflow orchestration service coordinating multiple Lambdas/services via a state machine, handling retries, parallelism, branching, and long-running flows (up to 1 year). Standard suits long flows requiring exactly-once execution; Express suits high-frequency, short flows. Exam trap: 'orchestrate multi-step flows, Lambda chains over 15 minutes, need visual retry logic' β†’ Step Functions.

Kinesis: Data Streams suits real-time streaming with custom logic β€” sub-second latency, ordered and replayable, scaled by manual sharding (or on-demand mode). Data Firehose delivers stream data to S3/Redshift/OpenSearch β€” near-real-time, no replay, fully managed automatic scaling. Managed Service for Apache Flink (formerly Analytics) does real-time SQL analytics on streams, also auto-scaling. Key numbers: Data Streams writes 1MB/s per shard (~1,000 records/s), reads 2MB/s; max record blob is 1MB; default retention 24 hours (extendable to 365 days); Firehose's minimum buffer is 60 seconds. Exam trap: 'just deliver stream data to S3/Redshift, no ops' β†’ Firehose; 'custom real-time processing, multiple consumers, replay needed' β†’ Data Streams; 'real-time SQL analytics on streams' β†’ Managed Service for Apache Flink. When ordering matters, use Data Streams with a partition key.

7. Monitoring & Management

CloudWatch: covers Metrics, Logs, Alarms, Dashboards, and Events/EventBridge. Standard metrics collect every 5 minutes, detailed monitoring every 1 minute, custom metrics down to 1 second (high resolution). OS-level metrics like memory/disk require installing the CloudWatch Agent. Exam trap: EC2 memory/disk usage is not a default metric β€” it requires the CloudWatch Agent. Alarms can trigger Auto Scaling, SNS notifications, or EC2-related actions.

CloudTrail: Management vs Data Events: Management Events are control-plane operations (creating/deleting resources, sign-ins), on by default and free. Data Events are high-frequency data-plane operations (S3 object-level GET/PUT, Lambda invocations), off by default, must be separately enabled, and cost extra. Config: records resource configuration change history and evaluates compliance via Config Rules (e.g., 'all EBS volumes must be encrypted'); Conformance Packs bundle-deploy rule sets, with SSM Automation enabling auto-remediation. Parameter Store vs Secrets Manager: Parameter Store is for config/simple secrets, free at the standard tier but with no built-in auto-rotation (build your own), size limits of 4KB standard/8KB advanced. Secrets Manager is dedicated secret management with built-in auto-rotation (e.g., RDS credentials), charged per secret per month, with a 64KB size limit. Exam trap: 'need auto-rotation of database credentials' β†’ Secrets Manager; 'simple config/parameter storage, cost-sensitive' β†’ Parameter Store. Other management tools: Trusted Advisor checks best practices for cost/performance/security/fault-tolerance/service limits (full checks require Business/Enterprise Support). Cost Explorer visualizes historical cost and forecasts. Budgets sets threshold alerts (cost, usage, RI/SP coverage). Compute Optimizer uses ML to recommend optimal EC2/ASG/Lambda/EBS sizing for cost savings.

8. Migration & Transfer

Snow bandwidth math: the core formula is 'transfer time = data size Γ· available bandwidth.' Rule of thumb: if online transfer would exceed roughly a week, consider a physical device. Quick conversions (bytes): 100Mbps β‰ˆ 12.5MB/s; 1Gbps β‰ˆ 125MB/s; 10Gbps β‰ˆ 1.25GB/s. For example, transferring 100TB at full 1Gbps takes roughly 9–12 days theoretically; at only 100Mbps it takes roughly 3 months β€” at that point a Snow physical device should be considered (though DataSync or the newer Snowball Edge is the real-world preference). Reference: 100Mbpsβ‰ˆ12.5MB/s, 1Gbpsβ‰ˆ125MB/s, 10Gbpsβ‰ˆ1.25GB/s; 10TB over 100Mbps takes roughly 9+ days.

DataSync vs Transfer Family vs Storage Gateway: DataSync is an online, agent-based method for bulk migrating files/objects (NFS/SMB/HDFS β†’ S3/EFS/FSx), offering acceleration, verification, and scheduling (as fast as hourly), non-real-time. Transfer Family provides managed SFTP/FTPS/FTP endpoints landing files directly in S3/EFS, for partners exchanging data continuously over standard protocols. Storage Gateway provides hybrid cloud capability for ongoing on-premises access to cloud storage (see section 2.5). Exam trap: 'one-time or periodic bulk file migration to AWS' β†’ DataSync; 'partners uploading continuously via SFTP' β†’ Transfer Family; 'on-prem app needs continuous low-latency cloud access' β†’ Storage Gateway.

DMS: homogeneous vs heterogeneous: DMS migrates databases to AWS, supporting CDC-based continuous replication for minimal downtime while the source stays online. Homogeneous migration is same-engine (Oracle→Oracle, MySQL→RDS MySQL), needing no conversion. Heterogeneous migration is cross-engine (Oracle→Aurora PostgreSQL, SQL Server→MySQL), requiring SCT (Schema Conversion Tool) to convert schema/stored procedures/views first, then DMS to migrate the data itself. Exam trap: heterogeneous migrations (e.g., Oracle→PostgreSQL) require SCT for schema conversion — DMS itself only moves data. For minimal-downtime migration, use DMS + CDC. MGN (Application Migration Service): AWS's primary lift-and-shift/rehost service, using a lightweight agent for block-level continuous replication of physical/virtual/cloud servers into EC2, with a short downtime window (minutes), no application changes needed. It replaces the retired SMS and CloudEndure, sharing an engine with the DR service DRS, but MGN is for one-time migrations. Exam trap: 'migrate whole servers to EC2, no app changes, minimal downtime' → MGN; 'cross-engine database migration' → DMS+SCT; 'file migration' → DataSync.

9. Cost Optimization

Reserved Instances: Standard RI offers up to ~72% off, locking instance family/region/OS, but AZ/instance size (same family)/scope can change, and it can be resold on the marketplace β€” suited for stable, predictable workloads. Convertible RI offers ~54–66% off, letting you change instance family/OS/term β€” suited for long-term but changing needs. 1-year vs 3-year: 3-year commitments give deeper discounts but longer lock-in. Payment options: All/Partial/No Upfront. RIs cover EC2/RDS/Redshift/ElastiCache/OpenSearch (not Fargate/Lambda).

Savings Plans: commit to an hourly spend amount (1 or 3 years). Compute Savings Plans are the most flexible (up to ~66% off, on par with Convertible RI), automatically covering EC2/Fargate/Lambda across regions and instance families. EC2 Instance Savings Plans lock the instance family/region but offer the highest discount (~72%, on par with Standard RI). Since 2025, Savings Plans are recommended over RIs in most scenarios for their flexibility and lower operational overhead. Exam trap: 'workload may move to Fargate/Lambda or another region, want savings plus flexibility' β†’ Compute Savings Plans (not the locked Standard RI); 'need to reserve capacity in a specific AZ' β†’ Standard RI (Savings Plans don't reserve capacity themselves).

Spot Instances: up to ~90% off, potentially interrupted with just 2 minutes' notice. Interruption handling includes checkpointing, decoupling via SQS, and stateless design. Spot Fleet/EC2 Fleet mix Spot and On-Demand across instance types/AZs to maintain target capacity, suited for batch processing, CI/CD, big data, and fault-tolerant web tiers. S3 cost optimization patterns: Lifecycle transitions to IA/Glacier; Intelligent-Tiering for unknown access patterns; S3 Storage Lens for analysis; merging small files to avoid Glacier's 40KB overhead; S3 Bucket Keys to lower KMS costs; and cleaning up old versions/incomplete multipart uploads. Compute Optimizer: uses ML on historical usage to recommend optimal EC2/ASG/Lambda/EBS sizing (right-sizing) for cost savings.

10. High Availability & DR

RPO vs RTO: RPO (Recovery Point Objective) is the maximum tolerable data loss, measured in time (e.g., 'at most 5 minutes of data loss' means backing up/replicating every 5 minutes). RTO (Recovery Time Objective) is the maximum tolerable downtime (e.g., 'must recover within 1 hour'). Smaller values for both mean higher cost.

DR strategies (cost/RTO from lowest to highest): Backup & Restore has RPO/RTO in hours, backing up to S3/another region and rebuilding on disaster, lowest cost. Pilot Light has RPO/RTO in minutes-to-hours, with core components (e.g., the database) running continuously while the rest stays off, spun up and scaled on disaster, low-moderate cost. Warm Standby has RPO/RTO in seconds-to-minutes, a scaled-down full-featured environment running continuously, just scaling up on disaster, moderate cost. Multi-Site Active-Active has near-zero RPO, with multiple regions serving at full scale simultaneously, highest cost. The key Pilot Light vs Warm Standby difference: Pilot Light cannot directly serve requests on disaster β€” servers must first be 'turned on' or deployed and scaled. Warm Standby is already running (at reduced capacity) and can accept traffic immediately, just needing to scale up.

Exam trap: 'lowest cost, hours-level recovery acceptable' β†’ Backup & Restore; 'core database continuously replicated but compute off, moderate cost' β†’ Pilot Light; 'scaled-down environment always running, quick to scale, RTO in minutes' β†’ Warm Standby; 'zero downtime, multiple regions active simultaneously' β†’ Multi-Site Active-Active. Pilot Light and Warm Standby are the most confusable pair β€” the test is whether traffic can be served immediately: Warm Standby can, Pilot Light can't. Multi-AZ vs Multi-Region: Multi-AZ spans multiple AZs within the same region, protecting against AZ-level failures (HA) via synchronous replication with low latency. Multi-Region spans multiple regions, protecting against region-level disasters, meeting compliance data-residency needs, and enabling global low latency β€” but is more complex and costly. Exam trap: ordinary HA needs are met by Multi-AZ alone; only choose Multi-Region for 'surviving an entire region outage' or 'compliance requires multiple regions.'

Back up stateless EC2 via AMI, not EBS snapshots: For stateless EC2 instances, back up/restore using AMIs (which include launch config and can include data volumes) β€” combined with a Launch Template and ASG for fast rebuilding. Data itself should be externalized to RDS/DynamoDB/S3/EFS. EBS snapshots only back up a single data volume, suited to stateful data disks. Exam trap: 'quickly rebuild a stateless web tier' β†’ AMI + Launch Template + ASG; 'back up a stateful data volume' β†’ EBS snapshot. RDS PITR vs manual snapshot: PITR is based on automated backups plus transaction logs, restoring to any second within the retention window (1–35 days), always restoring to a new instance. A manual snapshot is a user-triggered full backup at a point in time, kept permanently (surviving DB deletion), copyable cross-region/cross-account, suited for long-term archiving or migration. Exam trap: PITR only works within the retention window (default 7 days); for long-term/permanent retention or migration, use a manual snapshot. Either way, restoring always creates a new instance (never overwrites the original).

Appendix: Global Quick-Reference for Common Choices

Below is the highest-frequency mapping table across the whole cheat sheet β€” memorize it in reverse so a keyword instantly maps to a service name. Microsecond-level caching β†’ DAX; millisecond-level caching β†’ ElastiCache. Static IP + non-HTTP protocol + fast failover β†’ Global Accelerator; caching web content β†’ CloudFront. Blocking a specific IP β†’ NACL Deny rule; referencing security groups or instance-level control β†’ Security Group. Overlapping-CIDR connectivity β†’ PrivateLink; many interconnected VPCs β†’ Transit Gateway; two directly connected VPCs β†’ VPC Peering. Auto-rotating secrets β†’ Secrets Manager; simple config storage β†’ Parameter Store. Strictly ordered messages β†’ SQS FIFO; fan-out delivery β†’ SNS; event routing β†’ EventBridge. Cross-engine database migration β†’ DMS+SCT; whole-server migration β†’ MGN; file migration β†’ DataSync. Authenticating and issuing JWTs β†’ Cognito User Pool; exchanging for temporary AWS credentials β†’ Cognito Identity Pool. 'Who changed a resource' β†’ CloudTrail; 'is configuration compliant' β†’ Config; 'is there a threat' β†’ GuardDuty; 'is there a vulnerability' β†’ Inspector; 'is there sensitive data in S3' β†’ Macie. Cross-engine migration with minimal downtime β†’ DMS (with CDC); data-warehouse analytics β†’ Redshift; direct SQL on S3 data β†’ Athena.

Recommendations

First, memorize the number table thoroughly β€” every πŸ”‘ boxed number in this sheet is an instant-answer point on multiple-choice questions (15 minutes, 400KB, 3000/1000, 1–35 days, eleven nines, 180 days, gp3's 3000+125, etc.), and these carry high point value with fast decision speed. Second, focus on five key comparison groups: Security Groups vs NACL, Multi-AZ vs Read Replica, the five security services (CloudTrail/Config/GuardDuty/Inspector/Macie), SQS Standard vs FIFO, and the four DR strategies β€” these make up a large share of SAA questions, so be able to map keywords in the prompt straight to the answer. Third, use decision trees when answering β€” find the keyword first (static IP, microsecond, CIDR overlap, strict ordering, cross-region compliance) and map it directly to a service, avoiding time wasted agonizing between similar services. Fourth, know the gap between reality and the question bank β€” the Snow Family is largely retired but the question bank may still use old figures, where the traditional answer still applies; in real work, use DataSync instead. Similarly, CloudEndure/SMS have been replaced by MGN. Fifth, learn the thresholds that flip the answer β€” for RTO in minutes choose Warm Standby, in hours choose Pilot Light, and for longer acceptable windows choose Backup & Restore; when bandwidth times time exceeds a week, consider a physical migration device; for SSE-KMS throttling under high concurrency, enable S3 Bucket Keys; copying an encrypted snapshot cross-region requires switching to the destination region's KMS key. Sixth, get hands-on β€” do at least one console exercise each for Lambda concurrency settings, S3 lifecycle rules, RDS snapshot restore, and Security Group/NACL rules, which cements memory far better than reading alone.

Caveats

Numbers keep changing: AWS service limits are continuously updated (e.g., Lambda's /tmp was raised from 512MB to up to 10GB, KMS rotation changed from 3 years to 1 year, and the SQS FIFO in-flight message limit rose from 20,000 to 120,000 in November 2024). This sheet was compiled from public materials and official documentation as of July 2026; on exam day, defer to the then-current official docs. On the Snow Family, the 80TB/100PB figures in exam scenarios are historical; as of November 2025 only the 210TB Snowball Edge remains and it no longer accepts new customers, while Snowmobile (March 2024) and Snowcone (November 2024) have both been retired. Note the distinction between soft and hard limits: most concurrency/quota numbers are soft limits that can be raised on request, but hard limits (Lambda's 15 minutes, 10GB memory, 6MB sync payload; DynamoDB's 400KB item; SQS's 256KB message) cannot be changed. On furigana scope, common everyday kanji in Japanese are annotated with readings, while AWS proper nouns (Lambda, DynamoDB, etc.) are left in English without annotation. On sourcing, key numbers have been verified against official AWS documentation or blogs (Transit Gateway pricing at $0.05/hr + $0.02/GB, Aurora Global Database's 'latency typically under a second,' Shield Advanced at $3,000/month, gp3 being ~20% cheaper than gp2, Savings Plans discounts up to 66%/72%, SQS in-flight messages around 120,000 β€” all confirmed). Some cost percentages and retrieval-time figures come from vendor blogs and are adopted where consistent with official AWS docs; a few individual 2025–2026 developments (such as recent naming-related news around MGN) are recent developments rather than core exam content, and can simply be remembered under the traditional name MGN.