SOA-C03 sample questions with answers

10 free practice questions for the AWS Certified CloudOps Engineer – Associate exam. Try each one, then open the answer to see why the right option wins and every other option loses.

Question 1Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A Lambda function returns correct results, but no log group appears for it in CloudWatch Logs, so the team cannot see any of its output. The function was deployed with a minimal execution role that a developer wrote by hand.

What is the cause?

  1. A.

    The function's reserved concurrency is set to zero, which suppresses log delivery to CloudWatch Logs while still allowing the function itself to run normally.

  2. B.

    CloudWatch Logs creates a log group for a function only after it has been invoked at least one hundred times, and this function has not reached that count.

  3. C.

    Lambda writes output to CloudWatch Logs only when the function's logging configuration sets the system log level to DEBUG rather than the default INFO level.

  4. D.

    The execution role does not grant logs:CreateLogGroup, logs:CreateLogStream and logs:PutLogEvents.

Show answer

Answer: D

Lambda writes logs using the function's execution role, so a role without CloudWatch Logs permissions produces no log group at all.

  • A. Reserved concurrency of zero would stop the function from running entirely; it does not selectively disable logging.
  • B. Lambda creates the log group on the first invocation that has permission to create it; there is no invocation-count threshold.
  • C. Log level controls which platform events are emitted, not whether the function's own output reaches CloudWatch Logs at all.
  • D. Correct. The runtime writes logs with the execution role's credentials, so missing logs permissions means no log group is ever created.
Question 2Monitoring, Logging, Analysis, Remediation, and Performance Optimization

An application publishes a custom heartbeat metric every minute while it is healthy and publishes nothing at all when it crashes. An alarm on that metric uses a 60-second period, 2 out of 2 datapoints to alarm, and the missing data treatment set to ignore.

What happens when the application crashes?

  1. A.

    The alarm moves to ALARM after two minutes, because CloudWatch treats an absent datapoint as a datapoint that failed the threshold comparison.

  2. B.

    The alarm moves to INSUFFICIENT_DATA after two minutes and invokes its insufficient-data action if one is configured, but it never reaches ALARM.

  3. C.

    The alarm keeps the OK state it held before the crash and never transitions, so nobody is notified. The treatment should be breaching.

  4. D.

    The alarm repeats the last datapoint it received for up to three hours, then moves to INSUFFICIENT_DATA once that replay window is exhausted.

Show answer

Answer: C

The ignore treatment freezes the alarm's current state, so a metric that simply stops publishing never raises an alert.

  • A. Treating absent datapoints as breaching is the breaching treatment, which is precisely the setting this alarm does not use.
  • B. INSUFFICIENT_DATA is what the missing treatment produces when the whole evaluation range is empty, not what ignore produces.
  • C. Correct. ignore maintains the current state, so an alarm sitting in OK stays in OK however long the metric is absent.
  • D. CloudWatch never replays or repeats the last datapoint; there is no such evaluation behaviour.
Question 3Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A company runs 2,000 EC2 instances with the CloudWatch agent collecting eight metrics each. The agent configuration appends the InstanceId dimension to every metric and also sets aggregation_dimensions to roll the metrics up by AutoScalingGroupName. The custom-metric charge is far higher than expected, and every dashboard and alarm the team uses reads only the per-Auto Scaling group rollup.

What should the engineer change?

  1. A.

    Remove InstanceId from append_dimensions so that only the AutoScalingGroupName rollup is published, because every unique dimension combination is billed as its own custom metric and 2,000 instances times eight metrics is 16,000 of them.

  2. B.

    Write the metrics into the AWS/EC2 namespace instead of CWAgent, because metrics published to an AWS-reserved namespace are not charged at the custom-metric rate.

  3. C.

    Raise metricscollectioninterval from 60 seconds to 300 seconds, because custom metrics are billed per datapoint published rather than per unique metric name and dimension set.

  4. D.

    Replace the agent's publishing with a CloudWatch metric stream to Amazon Data Firehose, because streamed metrics are billed per GB delivered rather than per metric.

Show answer

Answer: A

Custom metrics are billed per unique name and dimension combination, so the per-instance dimension is what makes 16,000 metrics.

  • A. Correct. Dimension cardinality drives custom-metric count, and the per-instance series is not being used.
  • B. AWS-reserved namespaces such as AWS/EC2 cannot be written to by customers, and namespace choice does not change custom-metric pricing.
  • C. Custom metrics are charged per unique metric per month; interval affects request volume, not the metric count.
  • D. A metric stream is an additional export of existing metrics, so it adds cost rather than removing the per-instance metrics.
Question 4Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A downstream order-fulfilment service was unavailable for 36 hours. Events published to a custom event bus during that window reached a dead-letter queue for one target and were lost for two targets that had no dead-letter queue. The team must be able to reprocess events after any future outage of this length.

Which two actions should they take? (Choose TWO.)

Choose 2.

  1. A.

    Attach a dead-letter queue to every target and write a Lambda function that republishes messages from those queues back onto the event bus.

  2. B.

    Set every target's retry policy to a maximum event age of seven days so that EventBridge holds undelivered events until the target recovers.

  3. C.

    Create an EventBridge archive on the custom event bus with a retention period covering the longest outage they intend to recover from.

  4. D.

    After an outage, start a replay from the archive for the affected time window, selecting only the rules whose targets must be reprocessed.

  5. E.

    Enable an event replication configuration on the bus so that events are copied to a second Region and can be replayed from that Region instead.

Show answer

Answer: C, D

An archive captures events as they arrive on the bus, and a replay re-delivers a chosen time window to chosen rules.

  • A. This recovers only what a dead-letter queue captured, which excludes the two targets that had none.
  • B. The maximum event age on an EventBridge retry policy caps at 24 hours, which cannot cover a 36-hour outage.
  • C. Correct. An archive on the bus captures events independently of any target's delivery outcome.
  • D. Correct. A replay re-delivers a chosen time window to chosen rules, so only the affected targets reprocess.
  • E. Cross-Region replication copies events elsewhere; it does not create the time-window replay capability the team needs.
Question 5Monitoring, Logging, Analysis, Remediation, and Performance Optimization

An operations team has built a closed loop: a CloudWatch alarm on a custom queue-depth metric, an EventBridge rule that matches the alarm's state change to ALARM, and a Systems Manager Automation runbook, with its service role, that restarts the stuck worker. Before relying on it, the team wants one test that proves the whole chain works.

Which test should the team perform?

  1. A.

    Test the rule's event pattern in the EventBridge sandbox against a sample alarm state change event and confirm that the pattern matches.

  2. B.

    Start the Automation runbook manually from the Systems Manager console with the stuck worker's instance ID as input and confirm that the execution completes successfully.

  3. C.

    Use the SetAlarmState API to put the alarm into ALARM, then confirm in the Automation execution history that the runbook was started by the rule and completed.

  4. D.

    Add the runbook's ARN to the alarm's ALARM actions as well, then wait for the next real backlog to confirm that the worker restarts.

Show answer

Answer: C

SetAlarmState emits a genuine alarm state change event, so it tests the EventBridge rule, its permissions and the runbook in one step.

  • A. The sandbox validates pattern syntax and matching only; it does not deliver the event, so the role and the runbook are not tested.
  • B. This proves only the runbook; the alarm, the event pattern and the rule's permission to start the automation are never exercised.
  • C. A forced state change emits a real alarm state change event, so the test exercises the rule's pattern, its role and the runbook end to end.
  • D. An Automation runbook is not a supported alarm action, and waiting for a real incident is not a test performed before relying on the loop.
Question 6Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A platform team must publish the same eight-widget CloudWatch dashboard into 40 member accounts across three Regions, so that each account's operators open it in their own account. The widget set changes every few weeks and must stay identical everywhere, and new accounts join the organisational unit every month.

Which two actions meet this with the LEAST ongoing effort? (Choose TWO.)

Choose 2.

  1. A.

    Define the dashboard as an AWS::CloudWatch::Dashboard resource whose DashboardBody is rendered from template and pseudo parameters.

  2. B.

    Deploy the template with CloudFormation StackSets using service-managed permissions, targeting the organisational unit and the three Regions with automatic deployment enabled.

  3. C.

    Create the dashboard by hand in one account, export its source JSON from the console, and paste it into the other 39 accounts on every change.

  4. D.

    Write a Lambda function in a central account that assumes a role in every member account each night and calls PutDashboard with a rendered body.

  5. E.

    Build the dashboard once in a monitoring account with cross-account observability enabled and share it read-only with every member account.

Show answer

Answer: A, B

A dashboard is a CloudFormation resource, and StackSets distributes and updates it across every account and Region from one definition.

  • A. Correct. The dashboard body belongs in a CloudFormation template so that one definition produces every copy.
  • B. Correct. StackSets with service-managed permissions and automatic deployment covers current and future accounts in the OU.
  • C. Manual copying across 40 accounts and three Regions drifts immediately and repeats in full on every change.
  • D. This works but rebuilds cross-account deployment, drift handling and retries that StackSets already provides.
  • E. Dashboard sharing exposes one dashboard to viewers; it does not create a dashboard inside each member account.
Question 7Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A security group that protected a payment service's database tier was deleted two days ago, and the service was briefly unreachable. The account has never had a CloudTrail trail or an event data store configured. The incident manager needs the IAM principal and the source IP address that made the deletion, and wants the answer from existing records rather than new tooling.

Which action will provide this information?

  1. A.

    In the AWS CloudTrail console, open Event history, filter on the event name DeleteSecurityGroup, and read the userIdentity and sourceIPAddress fields of the event.

  2. B.

    Query the VPC Flow Logs for the database subnets for the time of the deletion and identify the source IP address that sent the request to the security group.

  3. C.

    Run a CloudWatch Logs Insights query across the payment service's application log groups to find the request that deleted the security group.

  4. D.

    Open the AWS Config configuration history for the security group and read the principal and source IP address recorded in the configuration item for the deletion.

Show answer

Answer: A

CloudTrail Event history records the last 90 days of management events without a trail, including the userIdentity and sourceIPAddress for DeleteSecurityGroup.

  • A. Event history keeps 90 days of management events with no trail required, and each event records the caller identity and source IP address.
  • B. Flow logs record network traffic through network interfaces; a DeleteSecurityGroup call is an API request to the EC2 endpoint, which flow logs do not capture.
  • C. Application logs record what the application did, not control-plane API calls made against the account, so the deletion would not appear there.
  • D. A Config configuration item records the resource's state and that it was deleted, not the caller identity or source IP of the API request.
Question 8Monitoring, Logging, Analysis, Remediation, and Performance Optimization

Four EC2 instances of the same type run a latency-sensitive simulation in one Availability Zone. They were launched separately over several weeks, and inter-node network latency is now the limiting factor. The team wants the instances in a new cluster placement group and has a maintenance window in which they may be unavailable.

Which sequence of steps should the team follow?

  1. A.

    Create a cluster placement group in the instances' Availability Zone, modify each running instance's placement to the group, then reboot the instances so that the change takes effect.

  2. B.

    Create a spread placement group in the instances' Availability Zone, stop the instances, modify each instance's placement to the group, then start them and confirm.

  3. C.

    Create a cluster placement group in the instances' Availability Zone, stop the instances, modify each instance's placement to the group, then start them and confirm the placement.

  4. D.

    Create a cluster placement group in a different Availability Zone with more capacity, stop the instances, modify each instance's placement to the group, then start them.

Show answer

Answer: C

Create the cluster group in the same zone, stop the instances, modify their placement, and start them again; placement cannot change while running.

  • A. Placement cannot be changed on a running instance, and a reboot does not move an instance to a different host.
  • B. A spread group puts instances on distinct hardware to reduce correlated failure, which does not lower inter-node latency.
  • C. An instance must be stopped before its placement can be changed, and the cluster group must be in the zone the instances occupy.
  • D. Stopping and starting does not move an instance to another zone, and a cluster group is confined to one zone, so the change cannot succeed.
Question 9Monitoring, Logging, Analysis, Remediation, and Performance Optimization

An Auto Scaling group behind an Application Load Balancer scales on average CPU. Instances take about six minutes to boot, register and warm their cache. During a traffic ramp the group repeatedly adds instances, then removes them two minutes later, then adds them again.

What is the most likely cause?

  1. A.

    The scaling policy is a step scaling policy whose steps overlap, so more than one step is applied for the same alarm breach during each evaluation period.

  2. B.

    The load balancer's target group uses least outstanding requests routing, which sends a disproportionate share of traffic to the newest targets and drives their CPU up.

  3. C.

    New instances report high CPU while they warm up and are still counted in the group's aggregated metric, so the policy over-scales and then scales back in once they settle.

  4. D.

    The group's termination policy is set to OldestInstance, so each scale-in removes an established instance and leaves only recently launched instances to serve the traffic.

Show answer

Answer: C

Instances that are still warming up publish unrepresentative metrics, and including them in the group average makes the policy oscillate.

  • A. Step scaling applies the single step that matches the breach size; overlapping steps are not how the policy evaluates.
  • B. Routing algorithm changes the distribution of requests but does not produce a repeating scale-out and scale-in cycle.
  • C. Correct. Unwarmed instances skew the group's aggregated metric, causing the policy to over-scale and then reverse.
  • D. The termination policy decides which instance is removed during a scale-in; it does not cause the repeated add and remove cycle.
Question 10Monitoring, Logging, Analysis, Remediation, and Performance Optimization

A company tags all EC2 instances with a CostCenter tag. The analytics team's instances, tagged CostCenter=analytics, are t3.xlarge instances that run a message-parsing service 24 hours a day. CloudWatch shows CPUUtilization averaging 75% at all hours, CPUCreditBalance at 0, and a steadily growing CPUSurplusCreditsCharged metric. The monthly bill for these instances has risen well above the on-demand price of the instance type. AWS Compute Optimizer flags the instances as 'Not optimized'.

What should the operations engineer do to resolve the performance and cost issue?

  1. A.

    Move the instances into a cluster placement group so that they share CPU credits more efficiently.

  2. B.

    Enable EBS optimization on the instances so that storage I/O no longer competes with the CPU for credits.

  3. C.

    Migrate the instances to a fixed-performance instance family of similar size, such as m6i.xlarge, following the Compute Optimizer recommendation.

  4. D.

    Change the instances from unlimited mode to standard mode so that surplus credits are no longer charged and the instances throttle back to their baseline CPU performance whenever the credit balance is empty.

Show answer

Answer: C

A burstable instance sustaining 75% CPU around the clock burns surplus credits at a premium; a fixed-performance instance of equivalent size is both faster and cheaper, which is what Compute Optimizer is signalling.

  • A. Placement groups influence network locality only; CPU credits are per instance and cannot be pooled.
  • B. EBS optimisation provides dedicated storage throughput; it has no effect on CPU credit consumption.
  • C. Sustained high CPU on a burstable instance is cheaper and faster on a fixed-performance family; Compute Optimizer's recommendation reflects that.
  • D. Standard mode stops the surplus charges but throttles the instance to its 40% baseline, degrading a service that needs 75% CPU.

Keep going with 528 more SOA-C03 questions

Free papers every day, in the real exam formats, with progress by exam domain. Unlock every paper and timed mock exam when you are ready.

SOA-C03 sample questions with answers (10 free) · CertifyCloudx