TL;DR
Certification earned
I organized these study notes around the official SOA-C03 domains, highlighting core operational monitoring, automated systems management, reliability patterns, security compliance, networking troubleshooting, and practice scenarios.
Exam Overview and Domain Breakdown
The AWS Certified CloudOps Engineer – Associate (SOA-C03) exam covers five primary domains:
| Domain | Description | Exam Weighting |
|---|---|---|
| Domain 1 | Monitoring, Logging, Analysis, Remediation, and Performance Optimization | 22% |
| Domain 2 | Reliability and Business Continuity | 22% |
| Domain 3 | Deployment, Provisioning, and Automation | 22% |
| Domain 4 | Security and Compliance | 16% |
| Domain 5 | Networking and Content Delivery | 18% |
The exam consists of 65 questions (50 scored questions and 15 unscored pretest questions). The time limit is 130 minutes, and the passing score is 720 out of 1000 on a scaled score.
Here is the study sequence I followed:
Domain 1: Monitoring, Logging, Analysis, Remediation, and Performance Optimization
This domain focuses on configuring observability tools, analyzing metric spikes and status check failures, protecting log streams, and automating remediation workflows.
1. CloudWatch Metrics & OS-Level Agent
- Hypervisor vs. OS Metrics: Standard CloudWatch collects hypervisor-level metrics (CPU utilization, Disk I/O, Network in/out) out-of-the-box. OS-level metrics like Memory (RAM) utilization and Disk space usage require installing and configuring the Unified CloudWatch Agent.
- Status Checks Troubleshooting:
- System Status Check Failure: Indicates physical host hardware or network issues. Resolve by stopping and starting the EC2 instance to migrate it to a healthy physical host.
- Instance Status Check Failure: Indicates OS kernel corruption, memory exhaustion, or misconfigured software. Resolve by rebooting or inspecting system console logs.
- CloudWatch Logs Data Protection: Automatically scans and redacts protected health information (PHI) or personally identifiable information (PII) in log streams, emitting the
LogEventsWithFindingsmetric for security alarms.
2. Automated Remediation Workflows
- Event-Driven Remediation: Use Amazon EventBridge rules to match specific event patterns (such as
CloudWatch Alarm State Change) and trigger AWS Systems Manager Automation runbooks directly without maintaining custom Lambda code. - Lambda Performance Tuning: Fix cold start latency and high code initialization time by enabling Provisioned Concurrency to pre-warm function execution environments.
Domain 1 Practice Questions
Question 1
A medical company hosts an application on AWS. The company needs to be alerted if any logs in Amazon CloudWatch Logs contain protected health information (PHI). Which solution will meet this requirement?
- A. Configure Amazon Macie to process CloudWatch Logs and trigger an SNS alert
- B. Configure Amazon GuardDuty to process sensitive data findings
- C. Enable CloudWatch Logs data protection policies and create a CloudWatch alarm based on the
LogEventsWithFindingsmetric - D. Deploy an AWS Lambda function to poll log groups continuously and parse text strings for PHI patterns
Show answer and reason
Answer: C. Enable CloudWatch Logs data protection policies and create a CloudWatch alarm based on the LogEventsWithFindings metric.
Reason: CloudWatch Logs data protection policies natively audit log streams in real time for sensitive data (PII/PHI) and automatically emit the LogEventsWithFindings metric, which can invoke CloudWatch alarms directly. Macie only scans S3 buckets, and GuardDuty detects network threats, not log contents.
Question 2
A company created a serverless application based on AWS Lambda functions. Response times are higher than expected because one function spends extra time during code initialization. Which solution resolves this issue with the LEAST development effort?
- A. Modify the initialization code to run asynchronously
- B. Configure reserved concurrency on the Lambda function
- C. Configure provisioned concurrency on the Lambda function
- D. Create a second Lambda function to split the incoming invocations
Show answer and reason
Answer: C. Configure provisioned concurrency on the Lambda function.
Reason: Provisioned concurrency pre-warms function execution environments and executes initialization code ahead of time, eliminating cold start delays without requiring code changes.
Domain 2: Reliability and Business Continuity
This domain validates your ability to configure high availability, design disaster recovery strategies (RTO/RPO), enforce backup policies, and manage auto scaling lifecycles.
1. Disaster Recovery & Replication
- RTO vs. RPO:
- Recovery Time Objective (RTO): Acceptable duration of downtime before service restoration.
- Recovery Point Objective (RPO): Acceptable amount of data loss measured in time.
- S3 Cross-Region Replication (CRR) Prerequisites:
- S3 Versioning MUST be explicitly enabled on BOTH the source and destination buckets.
- Configure a replication rule on the source bucket and attach an IAM role granting S3 cross-region read/write permissions.
2. Auto Scaling & Load Balancing
- Auto Scaling Lifecycle Hooks:
- Use lifecycle hooks to put terminating instances into a
Terminating:Waitstate. - Gives custom scripts time to execute final actions (e.g., uploading data files from EBS to S3) before triggering
complete-lifecycle-actionto allow instance termination.
- Use lifecycle hooks to put terminating instances into a
- ALB vs. NLB Health Checks:
- Application Load Balancer (ALB): Layer 7 load balancer; uses HTTP/HTTPS path-based health checks (e.g.,
/healthwith status code 200–399). - Network Load Balancer (NLB): Layer 4 load balancer; handles millions of requests/sec using TCP/UDP port health checks.
- Application Load Balancer (ALB): Layer 7 load balancer; uses HTTP/HTTPS path-based health checks (e.g.,
Domain 2 Practice Questions
Question 1
A financial company must send a copy of all data on attached Amazon EBS volumes to an S3 bucket before an EC2 instance is terminated by an Auto Scaling group. In testing, instances terminate before the copy completes. How can a CloudOps engineer resolve this?
- A. Configure EC2 Auto Scaling lifecycle hooks to put instances into a
Terminating:Waitstatus, and send thecomplete-lifecycle-actioncommand when complete - B. Enable S3 Transfer Acceleration and use multipart uploads
- C. Set the
DisableApiTerminationproperty totrueon the EC2 instances inside CloudFormation - D. Set the
AutoEnableIOproperty totrueon the attached EBS volumes
Show answer and reason
Answer: A. Configure EC2 Auto Scaling lifecycle hooks to put instances into a Terminating:Wait status, and send the complete-lifecycle-action command when complete.
Reason: ASG Lifecycle Hooks pause instance decommissioning and hold the instance in a Terminating:Wait state, allowing cleanup or backup scripts to finish before signal completion allows termination to proceed.
Question 2
A global company needs to replicate all new and modified objects from an S3 bucket in us-east-1 to a bucket in eu-west-1. Which configuration sequence ensures successful replication?
- A. Enable versioning on the destination bucket only, and create an S3 replication rule
- B. Enable versioning on both the source and destination buckets, create a replication rule in the source account, and attach an IAM role with S3 replication permissions
- C. Configure a replication rule on the source bucket, create an IAM role, and enable versioning afterward
- D. Enable versioning on the source bucket only, and configure an IAM role
Show answer and reason
Answer: B. Enable versioning on both the source and destination buckets, create a replication rule in the source account, and attach an IAM role with S3 replication permissions.
Reason: S3 Cross-Region Replication strictly requires versioning to be active on both buckets before rule creation, alongside an IAM role granting replication permissions.
Domain 3: Deployment, Provisioning, and Automation
This domain covers managing fleet systems using Systems Manager, detecting CloudFormation drift, automating patch management, and performing container deployments.
1. AWS Systems Manager (SSM) Engine
| Feature | Primary Operational Purpose |
|---|---|
| Session Manager | Shell-level access to instances without open inbound SSH/RDP ports, public IPs, or bastion hosts. |
| Patch Manager | Automates OS patching using Patch Baselines (approved patches) and Maintenance Windows. |
| Run Command | Executes administrative commands or scripts safely across instances at scale. |
| State Manager | Enforces OS software/configuration compliance continuously. |
| Automation Runbooks | Orchestrates complex multi-account operational procedures, approvals, and remediation tasks. |
2. Infrastructure as Code & Drift Detection
- CloudFormation Drift Detection: Scans stack instances or StackSets across accounts/regions to identify out-of-band manual changes made to deployed resources without updating templates.
- ECS Container Deployments:
- Rolling Deployment: Gradually replaces old tasks with new revisions behind an ALB. Cost-effective and maintains availability without extra target group charges.
- Blue/Green Deployment: Requires CodeDeploy and a secondary target group to shift traffic completely after validation.
Domain 3 Practice Questions
Question 1
A company needs to centralize daily security patching, health checks, and automated remediation across multiple AWS accounts and Regions, while supporting approval workflows for sensitive operations with the LEAST operational overhead. Which solution works best?
- A. Use Systems Manager State Manager associations to run command documents and Session Manager for manual remediation
- B. Configure AWS Config rules to trigger CloudFormation templates for patches and remediation
- C. Create Systems Manager Automation runbooks for patching and remediation, scheduled via State Manager associations with approvals configured
- D. Write custom Lambda functions orchestrated by Step Functions for patching and approval steps
Show answer and reason
Answer: C. Create Systems Manager Automation runbooks for patching and remediation, scheduled via State Manager associations with approvals configured.
Reason: Systems Manager Automation runbooks provide built-in multi-account patching, health checks, automated remediation, approval workflows, and centralized audit logging out-of-the-box with zero custom code maintenance.
Question 2
A company provisions resources across multiple AWS accounts using AWS CloudFormation StackSets. How can a CloudOps engineer detect manual changes made to these resources with the LEAST operational effort?
- A. Parse AWS CloudTrail logs using custom CloudWatch filter patterns
- B. Create CloudFormation hooks to validate resource configurations
- C. Run CloudFormation drift detection on all stack sets at regular intervals
- D. Deploy Service Control Policies (SCPs) to lock stack instances
Show answer and reason
Answer: C. Run CloudFormation drift detection on all stack sets at regular intervals.
Reason: StackSets Drift Detection natively compares deployed actual resource properties against the template baseline to highlight manual out-of-band modifications.
Domain 4: Security and Compliance
This domain covers identity federation, cross-account KMS access, governance guardrails using SCPs, and perimeter boundary protections.
1. Identity, Governance & KMS Security
- IAM OIDC Identity Providers: Connects external CI/CD tools (GitHub Actions, GitLab) to AWS using OpenID Connect federation. Grants temporary credentials via
AssumeRoleWithWebIdentitywithout static IAM access keys. - Cross-Account KMS Key Access: Requires permissions granted in BOTH the calling IAM policy AND the KMS Key Policy (resource policy) in the key’s home account.
- Service Control Policies (SCPs):
- Applied at the AWS Organizations level.
- An explicit
Denystatement in an SCP (e.g., denying actions outside approved regions likeus-east-1andeu-central-1) overrides all permissions across member accounts and cannot be bypassed.
2. Perimeter Security & Access Guardrails
- VPC Block Public Access (BPA): Account/VPC boundary guardrail that blocks inbound internet access across all VPC subnets in
ingress-onlymode while preserving outbound internet connectivity through NAT Gateways.
Domain 4 Practice Questions
Question 1
A CloudOps engineer must block inbound traffic from the internet to a VPC equipped with an attached Internet Gateway and NAT Gateways in public subnets with the LEAST operational overhead. Which solution meets this requirement?
- A. Set VPC Block Public Access (BPA) to
ingress-onlymode - B. Delete the default security group’s inbound
0.0.0.0/0rule - C. Replace the Internet Gateway with an Egress-Only Internet Gateway
- D. Add an explicit deny rule to the NAT Gateway security group
Show answer and reason
Answer: A. Set VPC Block Public Access (BPA) to ingress-only mode.
Reason: VPC Block Public Access in ingress-only mode blocks all incoming traffic from the public internet at the VPC boundary with zero subnet/routing changes, while preserving outbound traffic initiated via NAT Gateways.
Question 2
An IAM user in a QA account receives an AccessDenied error when attempting to encrypt an EBS volume using a KMS key located in a Development account. The QA user’s IAM policy correctly grants kms:Encrypt. How can the engineer resolve this?
- A. Attach
kms:Encryptpermissions to the root user in the QA account - B. Create a second KMS key inside the QA account
- C. Modify the resource policy (Key Policy) of the KMS key in the Development account to allow the QA user
- D. Run
aws configureon the QA user instance using the Development account access keys
Show answer and reason
Answer: C. Modify the resource policy (Key Policy) of the KMS key in the Development account to allow the QA user.
Reason: Cross-account KMS key access requires explicit permissions in both the requesting identity’s IAM policy AND the resource-based KMS Key Policy in the owner’s account.
Domain 5: Networking and Content Delivery
This domain covers VPC Peering route tables, S3 VPC Endpoints, Route 53 DNS verification, and Security Groups vs. NACLs.
1. VPC Endpoints & Route Tables
- Gateway VPC Endpoints (S3 & DynamoDB):
- Cost: Free to use (no hourly fees or per-GB data processing fees).
- Keeps data traffic internal to the AWS network without traversing the public internet.
- Interface VPC Endpoints (AWS PrivateLink): Attaches ENIs to subnets for most other AWS services; incurs hourly charges and data processing fees.
- VPC Peering Routing: Creating and accepting a peering connection is not enough. You MUST update route tables in BOTH VPCs to direct destination CIDRs to the
pcx-xxxpeering target.
2. Route 53 Mail Verification Records
- SPF Verification (TXT Records): To authorize Amazon SES to send emails on behalf of a custom domain and prevent recipient server rejections, publish a DNS TXT record containing SPF information:
v=spf1 include:amazonses.com -all
Domain 5 Practice Questions
Question 1
A batch job uploads 20 GB of data daily from EC2 instances in a private subnet to an Amazon S3 bucket. The engineer wants to eliminate data transfer and processing costs for future uploads. Which solution meets this requirement?
- A. Configure an Interface VPC Endpoint for Amazon S3
- B. Enable S3 Transfer Acceleration
- C. Configure an S3 File Gateway
- D. Configure a Gateway VPC Endpoint for Amazon S3
Show answer and reason
Answer: D. Configure a Gateway VPC Endpoint for Amazon S3.
Reason: Gateway VPC Endpoints for S3 and DynamoDB are completely free of charge (no hourly fees or data processing charges) and allow private connectivity directly from VPC subnets to S3.
Question 2
A CloudOps engineer configures a VPC peering connection between a Production VPC (10.0.0.0/16) and a Management VPC (172.16.0.0/16). The peering connection is active, and default NACLs are used, but instances cannot communicate across VPCs. How should this be resolved?
- A. Update the route tables in both VPCs to route target CIDR traffic through the peering connection
- B. Update security groups to enable cross-region VPC peering flags
- C. Change the Management VPC CIDR block to overlap with the Production VPC
- D. Replace default NACLs with custom allow rules
Show answer and reason
Answer: A. Update the route tables in both VPCs to route target CIDR traffic through the peering connection.
Reason: Establishing a VPC Peering connection creates the link, but traffic cannot pass until route table entries in both VPCs explicitly direct destination traffic to the pcx-xxx peering ID.
Five Scenario Questions for Final Review
Scenario 1: Preserving Local Data Prior to ASG Instance Teardown
An application running on EC2 instances behind an Auto Scaling group requires uploading temporary local transaction files to S3 before an instance terminates. Currently, instances terminate before the upload finishes. How should this be resolved?
- A. Enable S3 Transfer Acceleration on the target bucket
- B. Attach an ASG lifecycle hook for
EC2_INSTANCE_TERMINATINGto hold instances in aTerminating:Waitstate until the upload script completes - C. Modify the CloudFormation stack template to set
DisableApiTerminationtotrue - D. Increase the instance size to grant more CPU power for faster uploads
Show answer and reason
Answer: B. Attach an ASG lifecycle hook for EC2_INSTANCE_TERMINATING to hold instances in a Terminating:Wait state until the upload script completes.
Explanation: Lifecycle hooks pause instance shutdown, putting instances into a wait state where custom scripts can copy data to S3 before sending complete-lifecycle-action to resume termination.
Scenario 2: Automated Cost-Effective CPU Scaling
A web application on an Auto Scaling group behind an ALB experiences high CPU utilization during unpredictable peak hours, resulting in 503 errors. Which action resolves this in the MOST cost-effective way?
- A. Configure a target tracking scaling policy based on average CPU utilization
- B. Upgrade the launch template to a larger instance type
- C. Fix the Auto Scaling group size to maximum capacity permanently
- D. Configure the ALB to route traffic away from high CPU targets
Show answer and reason
Answer: A. Configure a target tracking scaling policy based on average CPU utilization.
Explanation: Target tracking policies scale capacity out during peak spikes and scale in during low-demand periods, ensuring capacity remains available while paying only for additional compute when needed.
Scenario 3: CI/CD Pipeline Without Static IAM Keys
A company wants to connect an external GitHub Actions CI/CD pipeline to AWS to deploy infrastructure automatically. To follow security best practices, no long-lived access keys can be used. What is the proper configuration?
- A. Create an IAM user and store the secret access keys inside GitHub Secrets
- B. Configure an IAM OpenID Connect (OIDC) Identity Provider, create an IAM role with web identity trust policies, and assume the role from GitHub Actions
- C. Configure IAM Identity Center to manage SAML single sign-on access for the pipeline
- D. Generate temporary credentials using AWS STS in a local terminal and commit them to the repository
Show answer and reason
Answer: B. Configure an IAM OpenID Connect (OIDC) Identity Provider, create an IAM role with web identity trust policies, and assume the role from GitHub Actions.
Explanation: IAM OIDC federation allows external automation tools to authenticate machine-to-machine, exchanging short-lived OIDC tokens for temporary AWS STS credentials without storing static IAM access keys.
Scenario 4: Isolating S3 Event Notification Failures
An event-driven architecture triggers Lambda functions whenever files arrive in S3. Files uploaded to a specific path (/orders/) fail to trigger Lambda, while files in all other folders trigger functions successfully. How should this be diagnosed?
- A. Increase the memory allocation and execution timeout of the Lambda function
- B. Verify and update the prefix filter configuration defined inside the S3 Event Notification settings
- C. Replace S3 Event Notifications with an S3 File Gateway
- D. Grant
s3:FullAccesspermissions to the Lambda execution role
Show answer and reason
Answer: B. Verify and update the prefix filter configuration defined inside the S3 Event Notification settings.
Explanation: S3 Event Notifications filter triggers using prefix and suffix rules. If notifications fail for one specific subfolder while succeeding everywhere else, the prefix filtering rule on the bucket event configuration is misconfigured.
Scenario 5: High-Throughput Layer 4 Health Checking
An architecture requires load balancing an API tier that processes millions of TCP requests per second across multiple Availability Zones. The health check must evaluate TCP port connectivity efficiently. Which configuration fulfills these requirements?
- A. Deploy an Application Load Balancer (ALB) with HTTP path health checks
- B. Deploy a Network Load Balancer (NLB) with TCP health checks on the target API port
- C. Deploy an ALB with TCP health checks on port 80
- D. Configure Route 53 latency routing directly to EC2 elastic IP addresses
Show answer and reason
Answer: B. Deploy a Network Load Balancer (NLB) with TCP health checks on the target API port.
Explanation: Network Load Balancers operate at Layer 4 (TCP/UDP) and are engineered to handle millions of requests per second with ultra-low latency while conducting native layer 4 TCP health checks.
Final Thoughts
Preparing for the CloudOps Engineer – Associate (SOA-C03) exam requires shifting from developer code implementation into infrastructure operations, VPC network routing, systems automation, and business continuity planning.
My advice