TL;DR

I prepared for SOA-C03 by mastering observability & alerting (CloudWatch Agent, Logs Insights, Data Protection), automated systems management (Systems Manager runbooks, Session Manager, Patch Manager), reliability & disaster recovery (RTO/RPO, AWS Backup, Auto Scaling lifecycle hooks), security guardrails (AWS Config, SCPs, KMS cross-account access), and VPC networking (Route tables, Security Groups vs. NACLs, Gateway VPC Endpoints). Combining operational automation with hands-on AWS administration patterns helped me pass the exam.

Certification earned

You can view my AWS Certified CloudOps Engineer – Associate (SOA-C03) certification on Credly.

I organized these study notes around the official SOA-C03 domains, highlighting core operational monitoring, automated systems management, reliability patterns, security compliance, networking troubleshooting, and practice scenarios.


Exam Overview and Domain Breakdown

The AWS Certified CloudOps Engineer – Associate (SOA-C03) exam covers five primary domains:

DomainDescriptionExam Weighting
Domain 1Monitoring, Logging, Analysis, Remediation, and Performance Optimization22%
Domain 2Reliability and Business Continuity22%
Domain 3Deployment, Provisioning, and Automation22%
Domain 4Security and Compliance16%
Domain 5Networking and Content Delivery18%

The exam consists of 65 questions (50 scored questions and 15 unscored pretest questions). The time limit is 130 minutes, and the passing score is 720 out of 1000 on a scaled score.

Here is the study sequence I followed:

flowchart TD A[Domain 1: Monitoring, Logging & Remediation] --> B[Domain 2: Reliability & Business Continuity] B --> C[Domain 3: Deployment & Automation] C --> D[Domain 4: Security & Compliance] D --> E[Domain 5: Networking & Content Delivery] E --> F[Scenario Review & Final Prep]

Domain 1: Monitoring, Logging, Analysis, Remediation, and Performance Optimization

This domain focuses on configuring observability tools, analyzing metric spikes and status check failures, protecting log streams, and automating remediation workflows.

1. CloudWatch Metrics & OS-Level Agent

  • Hypervisor vs. OS Metrics: Standard CloudWatch collects hypervisor-level metrics (CPU utilization, Disk I/O, Network in/out) out-of-the-box. OS-level metrics like Memory (RAM) utilization and Disk space usage require installing and configuring the Unified CloudWatch Agent.
  • Status Checks Troubleshooting:
    • System Status Check Failure: Indicates physical host hardware or network issues. Resolve by stopping and starting the EC2 instance to migrate it to a healthy physical host.
    • Instance Status Check Failure: Indicates OS kernel corruption, memory exhaustion, or misconfigured software. Resolve by rebooting or inspecting system console logs.
  • CloudWatch Logs Data Protection: Automatically scans and redacts protected health information (PHI) or personally identifiable information (PII) in log streams, emitting the LogEventsWithFindings metric for security alarms.

2. Automated Remediation Workflows

  • Event-Driven Remediation: Use Amazon EventBridge rules to match specific event patterns (such as CloudWatch Alarm State Change) and trigger AWS Systems Manager Automation runbooks directly without maintaining custom Lambda code.
  • Lambda Performance Tuning: Fix cold start latency and high code initialization time by enabling Provisioned Concurrency to pre-warm function execution environments.

Domain 1 Practice Questions

Question 1

A medical company hosts an application on AWS. The company needs to be alerted if any logs in Amazon CloudWatch Logs contain protected health information (PHI). Which solution will meet this requirement?

  • A. Configure Amazon Macie to process CloudWatch Logs and trigger an SNS alert
  • B. Configure Amazon GuardDuty to process sensitive data findings
  • C. Enable CloudWatch Logs data protection policies and create a CloudWatch alarm based on the LogEventsWithFindings metric
  • D. Deploy an AWS Lambda function to poll log groups continuously and parse text strings for PHI patterns
Show answer and reason

Answer: C. Enable CloudWatch Logs data protection policies and create a CloudWatch alarm based on the LogEventsWithFindings metric.

Reason: CloudWatch Logs data protection policies natively audit log streams in real time for sensitive data (PII/PHI) and automatically emit the LogEventsWithFindings metric, which can invoke CloudWatch alarms directly. Macie only scans S3 buckets, and GuardDuty detects network threats, not log contents.

Question 2

A company created a serverless application based on AWS Lambda functions. Response times are higher than expected because one function spends extra time during code initialization. Which solution resolves this issue with the LEAST development effort?

  • A. Modify the initialization code to run asynchronously
  • B. Configure reserved concurrency on the Lambda function
  • C. Configure provisioned concurrency on the Lambda function
  • D. Create a second Lambda function to split the incoming invocations
Show answer and reason

Answer: C. Configure provisioned concurrency on the Lambda function.

Reason: Provisioned concurrency pre-warms function execution environments and executes initialization code ahead of time, eliminating cold start delays without requiring code changes.


Domain 2: Reliability and Business Continuity

This domain validates your ability to configure high availability, design disaster recovery strategies (RTO/RPO), enforce backup policies, and manage auto scaling lifecycles.

1. Disaster Recovery & Replication

  • RTO vs. RPO:
    • Recovery Time Objective (RTO): Acceptable duration of downtime before service restoration.
    • Recovery Point Objective (RPO): Acceptable amount of data loss measured in time.
  • S3 Cross-Region Replication (CRR) Prerequisites:
    • S3 Versioning MUST be explicitly enabled on BOTH the source and destination buckets.
    • Configure a replication rule on the source bucket and attach an IAM role granting S3 cross-region read/write permissions.

2. Auto Scaling & Load Balancing

  • Auto Scaling Lifecycle Hooks:
    • Use lifecycle hooks to put terminating instances into a Terminating:Wait state.
    • Gives custom scripts time to execute final actions (e.g., uploading data files from EBS to S3) before triggering complete-lifecycle-action to allow instance termination.
  • ALB vs. NLB Health Checks:
    • Application Load Balancer (ALB): Layer 7 load balancer; uses HTTP/HTTPS path-based health checks (e.g., /health with status code 200–399).
    • Network Load Balancer (NLB): Layer 4 load balancer; handles millions of requests/sec using TCP/UDP port health checks.

Domain 2 Practice Questions

Question 1

A financial company must send a copy of all data on attached Amazon EBS volumes to an S3 bucket before an EC2 instance is terminated by an Auto Scaling group. In testing, instances terminate before the copy completes. How can a CloudOps engineer resolve this?

  • A. Configure EC2 Auto Scaling lifecycle hooks to put instances into a Terminating:Wait status, and send the complete-lifecycle-action command when complete
  • B. Enable S3 Transfer Acceleration and use multipart uploads
  • C. Set the DisableApiTermination property to true on the EC2 instances inside CloudFormation
  • D. Set the AutoEnableIO property to true on the attached EBS volumes
Show answer and reason

Answer: A. Configure EC2 Auto Scaling lifecycle hooks to put instances into a Terminating:Wait status, and send the complete-lifecycle-action command when complete.

Reason: ASG Lifecycle Hooks pause instance decommissioning and hold the instance in a Terminating:Wait state, allowing cleanup or backup scripts to finish before signal completion allows termination to proceed.

Question 2

A global company needs to replicate all new and modified objects from an S3 bucket in us-east-1 to a bucket in eu-west-1. Which configuration sequence ensures successful replication?

  • A. Enable versioning on the destination bucket only, and create an S3 replication rule
  • B. Enable versioning on both the source and destination buckets, create a replication rule in the source account, and attach an IAM role with S3 replication permissions
  • C. Configure a replication rule on the source bucket, create an IAM role, and enable versioning afterward
  • D. Enable versioning on the source bucket only, and configure an IAM role
Show answer and reason

Answer: B. Enable versioning on both the source and destination buckets, create a replication rule in the source account, and attach an IAM role with S3 replication permissions.

Reason: S3 Cross-Region Replication strictly requires versioning to be active on both buckets before rule creation, alongside an IAM role granting replication permissions.


Domain 3: Deployment, Provisioning, and Automation

This domain covers managing fleet systems using Systems Manager, detecting CloudFormation drift, automating patch management, and performing container deployments.

1. AWS Systems Manager (SSM) Engine

FeaturePrimary Operational Purpose
Session ManagerShell-level access to instances without open inbound SSH/RDP ports, public IPs, or bastion hosts.
Patch ManagerAutomates OS patching using Patch Baselines (approved patches) and Maintenance Windows.
Run CommandExecutes administrative commands or scripts safely across instances at scale.
State ManagerEnforces OS software/configuration compliance continuously.
Automation RunbooksOrchestrates complex multi-account operational procedures, approvals, and remediation tasks.

2. Infrastructure as Code & Drift Detection

  • CloudFormation Drift Detection: Scans stack instances or StackSets across accounts/regions to identify out-of-band manual changes made to deployed resources without updating templates.
  • ECS Container Deployments:
    • Rolling Deployment: Gradually replaces old tasks with new revisions behind an ALB. Cost-effective and maintains availability without extra target group charges.
    • Blue/Green Deployment: Requires CodeDeploy and a secondary target group to shift traffic completely after validation.

Domain 3 Practice Questions

Question 1

A company needs to centralize daily security patching, health checks, and automated remediation across multiple AWS accounts and Regions, while supporting approval workflows for sensitive operations with the LEAST operational overhead. Which solution works best?

  • A. Use Systems Manager State Manager associations to run command documents and Session Manager for manual remediation
  • B. Configure AWS Config rules to trigger CloudFormation templates for patches and remediation
  • C. Create Systems Manager Automation runbooks for patching and remediation, scheduled via State Manager associations with approvals configured
  • D. Write custom Lambda functions orchestrated by Step Functions for patching and approval steps
Show answer and reason

Answer: C. Create Systems Manager Automation runbooks for patching and remediation, scheduled via State Manager associations with approvals configured.

Reason: Systems Manager Automation runbooks provide built-in multi-account patching, health checks, automated remediation, approval workflows, and centralized audit logging out-of-the-box with zero custom code maintenance.

Question 2

A company provisions resources across multiple AWS accounts using AWS CloudFormation StackSets. How can a CloudOps engineer detect manual changes made to these resources with the LEAST operational effort?

  • A. Parse AWS CloudTrail logs using custom CloudWatch filter patterns
  • B. Create CloudFormation hooks to validate resource configurations
  • C. Run CloudFormation drift detection on all stack sets at regular intervals
  • D. Deploy Service Control Policies (SCPs) to lock stack instances
Show answer and reason

Answer: C. Run CloudFormation drift detection on all stack sets at regular intervals.

Reason: StackSets Drift Detection natively compares deployed actual resource properties against the template baseline to highlight manual out-of-band modifications.


Domain 4: Security and Compliance

This domain covers identity federation, cross-account KMS access, governance guardrails using SCPs, and perimeter boundary protections.

1. Identity, Governance & KMS Security

  • IAM OIDC Identity Providers: Connects external CI/CD tools (GitHub Actions, GitLab) to AWS using OpenID Connect federation. Grants temporary credentials via AssumeRoleWithWebIdentity without static IAM access keys.
  • Cross-Account KMS Key Access: Requires permissions granted in BOTH the calling IAM policy AND the KMS Key Policy (resource policy) in the key’s home account.
  • Service Control Policies (SCPs):
    • Applied at the AWS Organizations level.
    • An explicit Deny statement in an SCP (e.g., denying actions outside approved regions like us-east-1 and eu-central-1) overrides all permissions across member accounts and cannot be bypassed.

2. Perimeter Security & Access Guardrails

  • VPC Block Public Access (BPA): Account/VPC boundary guardrail that blocks inbound internet access across all VPC subnets in ingress-only mode while preserving outbound internet connectivity through NAT Gateways.

Domain 4 Practice Questions

Question 1

A CloudOps engineer must block inbound traffic from the internet to a VPC equipped with an attached Internet Gateway and NAT Gateways in public subnets with the LEAST operational overhead. Which solution meets this requirement?

  • A. Set VPC Block Public Access (BPA) to ingress-only mode
  • B. Delete the default security group’s inbound 0.0.0.0/0 rule
  • C. Replace the Internet Gateway with an Egress-Only Internet Gateway
  • D. Add an explicit deny rule to the NAT Gateway security group
Show answer and reason

Answer: A. Set VPC Block Public Access (BPA) to ingress-only mode.

Reason: VPC Block Public Access in ingress-only mode blocks all incoming traffic from the public internet at the VPC boundary with zero subnet/routing changes, while preserving outbound traffic initiated via NAT Gateways.

Question 2

An IAM user in a QA account receives an AccessDenied error when attempting to encrypt an EBS volume using a KMS key located in a Development account. The QA user’s IAM policy correctly grants kms:Encrypt. How can the engineer resolve this?

  • A. Attach kms:Encrypt permissions to the root user in the QA account
  • B. Create a second KMS key inside the QA account
  • C. Modify the resource policy (Key Policy) of the KMS key in the Development account to allow the QA user
  • D. Run aws configure on the QA user instance using the Development account access keys
Show answer and reason

Answer: C. Modify the resource policy (Key Policy) of the KMS key in the Development account to allow the QA user.

Reason: Cross-account KMS key access requires explicit permissions in both the requesting identity’s IAM policy AND the resource-based KMS Key Policy in the owner’s account.


Domain 5: Networking and Content Delivery

This domain covers VPC Peering route tables, S3 VPC Endpoints, Route 53 DNS verification, and Security Groups vs. NACLs.

1. VPC Endpoints & Route Tables

  • Gateway VPC Endpoints (S3 & DynamoDB):
    • Cost: Free to use (no hourly fees or per-GB data processing fees).
    • Keeps data traffic internal to the AWS network without traversing the public internet.
  • Interface VPC Endpoints (AWS PrivateLink): Attaches ENIs to subnets for most other AWS services; incurs hourly charges and data processing fees.
  • VPC Peering Routing: Creating and accepting a peering connection is not enough. You MUST update route tables in BOTH VPCs to direct destination CIDRs to the pcx-xxx peering target.

2. Route 53 Mail Verification Records

  • SPF Verification (TXT Records): To authorize Amazon SES to send emails on behalf of a custom domain and prevent recipient server rejections, publish a DNS TXT record containing SPF information:
    v=spf1 include:amazonses.com -all

Domain 5 Practice Questions

Question 1

A batch job uploads 20 GB of data daily from EC2 instances in a private subnet to an Amazon S3 bucket. The engineer wants to eliminate data transfer and processing costs for future uploads. Which solution meets this requirement?

  • A. Configure an Interface VPC Endpoint for Amazon S3
  • B. Enable S3 Transfer Acceleration
  • C. Configure an S3 File Gateway
  • D. Configure a Gateway VPC Endpoint for Amazon S3
Show answer and reason

Answer: D. Configure a Gateway VPC Endpoint for Amazon S3.

Reason: Gateway VPC Endpoints for S3 and DynamoDB are completely free of charge (no hourly fees or data processing charges) and allow private connectivity directly from VPC subnets to S3.

Question 2

A CloudOps engineer configures a VPC peering connection between a Production VPC (10.0.0.0/16) and a Management VPC (172.16.0.0/16). The peering connection is active, and default NACLs are used, but instances cannot communicate across VPCs. How should this be resolved?

  • A. Update the route tables in both VPCs to route target CIDR traffic through the peering connection
  • B. Update security groups to enable cross-region VPC peering flags
  • C. Change the Management VPC CIDR block to overlap with the Production VPC
  • D. Replace default NACLs with custom allow rules
Show answer and reason

Answer: A. Update the route tables in both VPCs to route target CIDR traffic through the peering connection.

Reason: Establishing a VPC Peering connection creates the link, but traffic cannot pass until route table entries in both VPCs explicitly direct destination traffic to the pcx-xxx peering ID.


Five Scenario Questions for Final Review

Scenario 1: Preserving Local Data Prior to ASG Instance Teardown

An application running on EC2 instances behind an Auto Scaling group requires uploading temporary local transaction files to S3 before an instance terminates. Currently, instances terminate before the upload finishes. How should this be resolved?

  • A. Enable S3 Transfer Acceleration on the target bucket
  • B. Attach an ASG lifecycle hook for EC2_INSTANCE_TERMINATING to hold instances in a Terminating:Wait state until the upload script completes
  • C. Modify the CloudFormation stack template to set DisableApiTermination to true
  • D. Increase the instance size to grant more CPU power for faster uploads
Show answer and reason

Answer: B. Attach an ASG lifecycle hook for EC2_INSTANCE_TERMINATING to hold instances in a Terminating:Wait state until the upload script completes.

Explanation: Lifecycle hooks pause instance shutdown, putting instances into a wait state where custom scripts can copy data to S3 before sending complete-lifecycle-action to resume termination.

Scenario 2: Automated Cost-Effective CPU Scaling

A web application on an Auto Scaling group behind an ALB experiences high CPU utilization during unpredictable peak hours, resulting in 503 errors. Which action resolves this in the MOST cost-effective way?

  • A. Configure a target tracking scaling policy based on average CPU utilization
  • B. Upgrade the launch template to a larger instance type
  • C. Fix the Auto Scaling group size to maximum capacity permanently
  • D. Configure the ALB to route traffic away from high CPU targets
Show answer and reason

Answer: A. Configure a target tracking scaling policy based on average CPU utilization.

Explanation: Target tracking policies scale capacity out during peak spikes and scale in during low-demand periods, ensuring capacity remains available while paying only for additional compute when needed.

Scenario 3: CI/CD Pipeline Without Static IAM Keys

A company wants to connect an external GitHub Actions CI/CD pipeline to AWS to deploy infrastructure automatically. To follow security best practices, no long-lived access keys can be used. What is the proper configuration?

  • A. Create an IAM user and store the secret access keys inside GitHub Secrets
  • B. Configure an IAM OpenID Connect (OIDC) Identity Provider, create an IAM role with web identity trust policies, and assume the role from GitHub Actions
  • C. Configure IAM Identity Center to manage SAML single sign-on access for the pipeline
  • D. Generate temporary credentials using AWS STS in a local terminal and commit them to the repository
Show answer and reason

Answer: B. Configure an IAM OpenID Connect (OIDC) Identity Provider, create an IAM role with web identity trust policies, and assume the role from GitHub Actions.

Explanation: IAM OIDC federation allows external automation tools to authenticate machine-to-machine, exchanging short-lived OIDC tokens for temporary AWS STS credentials without storing static IAM access keys.

Scenario 4: Isolating S3 Event Notification Failures

An event-driven architecture triggers Lambda functions whenever files arrive in S3. Files uploaded to a specific path (/orders/) fail to trigger Lambda, while files in all other folders trigger functions successfully. How should this be diagnosed?

  • A. Increase the memory allocation and execution timeout of the Lambda function
  • B. Verify and update the prefix filter configuration defined inside the S3 Event Notification settings
  • C. Replace S3 Event Notifications with an S3 File Gateway
  • D. Grant s3:FullAccess permissions to the Lambda execution role
Show answer and reason

Answer: B. Verify and update the prefix filter configuration defined inside the S3 Event Notification settings.

Explanation: S3 Event Notifications filter triggers using prefix and suffix rules. If notifications fail for one specific subfolder while succeeding everywhere else, the prefix filtering rule on the bucket event configuration is misconfigured.

Scenario 5: High-Throughput Layer 4 Health Checking

An architecture requires load balancing an API tier that processes millions of TCP requests per second across multiple Availability Zones. The health check must evaluate TCP port connectivity efficiently. Which configuration fulfills these requirements?

  • A. Deploy an Application Load Balancer (ALB) with HTTP path health checks
  • B. Deploy a Network Load Balancer (NLB) with TCP health checks on the target API port
  • C. Deploy an ALB with TCP health checks on port 80
  • D. Configure Route 53 latency routing directly to EC2 elastic IP addresses
Show answer and reason

Answer: B. Deploy a Network Load Balancer (NLB) with TCP health checks on the target API port.

Explanation: Network Load Balancers operate at Layer 4 (TCP/UDP) and are engineered to handle millions of requests per second with ultra-low latency while conducting native layer 4 TCP health checks.


Final Thoughts

Preparing for the CloudOps Engineer – Associate (SOA-C03) exam requires shifting from developer code implementation into infrastructure operations, VPC network routing, systems automation, and business continuity planning.

My advice

Focus heavily on understanding AWS Systems Manager (SSM) tools (Session Manager, Patch Manager, Runbooks), VPC route tables and endpoints, CloudWatch Agent OS metrics, S3 replication prerequisites, and Auto Scaling lifecycle hooks. Practice identifying cost-effective operational solutions with minimal maintenance overhead.