DEA-C01 sample questions with answers

10 free practice questions for the AWS Certified Data Engineer – Associate exam. Try each one, then open the answer to see why the right option wins and every other option loses.

Question 1Data Ingestion and Transformation

An ETL pipeline extracts from an operational database, applies three successive transformation stages, and loads the result into Amazon Redshift. Each stage is a separate AWS Glue job, and a failed stage must be re-runnable without re-extracting from the operational database.

Where should the output of each stage be written?

  1. A.

    To a stage-specific Amazon S3 prefix, so each job reads the previous stage's output and can be re-run independently

  2. B.

    To the local disk of the Glue workers, so intermediate results avoid Amazon S3 request charges and stay close to the compute

  3. C.

    To a Spark in-memory DataFrame passed directly between the three jobs, avoiding the cost of writing intermediate results to storage

  4. D.

    Back to the operational database in temporary tables, so all intermediate state is held in one transactional system

Show answer

Answer: A

Durable, addressable staging in Amazon S3 is what makes each stage independently re-runnable without touching the source.

  • A. Durable, addressable per-stage output survives the job, supports inspection, and makes each stage independently re-runnable.
  • B. Glue worker local disk is deallocated when the run ends, so the next stage finds nothing and a failed stage cannot resume.
  • C. Separate Glue jobs are separate Spark applications; in-memory DataFrames cannot cross between them and do not survive failure.
  • D. This puts ETL load and storage on the operational system and couples pipeline reliability to a transactional database.
Question 2Data Ingestion and Transformation

A company already runs an Amazon MSK cluster whose topics carry order events. A new requirement is to land every event in Amazon S3 as it arrives, with no application code to write or operate, reusing the open-source Kafka Connect ecosystem the team already knows.

Which approach should the company take?

  1. A.

    Create an Amazon Data Firehose stream that subscribes to the MSK topic as a consumer group and add a Lambda transformation that reformats each record

  2. B.

    Deploy an Amazon S3 sink connector on MSK Connect, which runs and scales Kafka Connect workers as a managed service

  3. C.

    Write a Managed Service for Apache Flink application that reads the topic and writes each record to Amazon S3

  4. D.

    Run a Kafka Connect cluster on Amazon EC2 instances in an Auto Scaling group, install the S3 sink connector plugin on each instance, and manage the worker configuration yourself

Show answer

Answer: B

MSK Connect is the managed Kafka Connect service, so an S3 sink connector meets the requirement with no code and no worker fleet to operate.

  • A. This adds an unnecessary Lambda transformation and abandons the Kafka Connect ecosystem the team wants to reuse.
  • B. MSK Connect runs and scales Kafka Connect workers as a managed service, so an S3 sink connector needs no code or fleet.
  • C. A Flink application is application code, which the requirement excludes, and it does not reuse Kafka Connect.
  • D. Self-managed Connect workers on EC2 restore all the operational work the managed service exists to remove.
Question 3Data Ingestion and Transformation

Event payloads arriving in Amazon Redshift contain a nested JSON attributes object whose keys differ from event to event. Analysts must query individual attributes with SQL without the team declaring a column for every possible key.

Which Amazon Redshift feature should the team use?

  1. A.

    A narrow table with one row per key-value pair, joined back to the fact table on every query

  2. B.

    A VARCHAR(65535) column holding the raw JSON, parsed in every query with a user-defined function written for each attribute

  3. C.

    A materialized view that flattens the JSON into a fixed set of columns, refreshed whenever a new key appears

  4. D.

    The SUPER data type, which stores semi-structured JSON and is queried with PartiQL path expressions

Show answer

Answer: D

The SUPER data type stores semi-structured data natively and is queried with PartiQL, so varying keys need no schema changes.

  • A. An entity-attribute-value table forces a join and aggregation on every query and discards type information.
  • B. This defers all parsing to query time and needs a new function per attribute, which is a schema change in disguise.
  • C. A flattened view fixes the columns at definition time, so an unanticipated key requires someone to redefine it.
  • D. SUPER stores hierarchical JSON natively and PartiQL path expressions reach any key without a DDL change.
Question 4Data Ingestion and Transformation

A team wants Amazon Data Firehose to convert incoming JSON records from a Kinesis data stream to Apache Parquet before writing them to Amazon S3.

What must be in place for Firehose record format conversion to work?

  1. A.

    A table in the AWS Glue Data Catalog that defines the target schema, which Firehose reads to perform the conversion

  2. B.

    An AWS Glue streaming ETL job subscribed to the Firehose stream that rewrites the delivered objects as Parquet

  3. C.

    A Firehose dynamic partitioning configuration, because format conversion is only available when dynamic partitioning is enabled

  4. D.

    An AWS Lambda transformation function on the Firehose stream that converts each record to Parquet before delivery, writing the columnar output to the destination bucket on the Firehose stream's behalf

Show answer

Answer: A

Firehose record format conversion reads the target schema from an AWS Glue Data Catalog table to serialize records as Parquet or ORC.

  • A. Format conversion serializes to Parquet or ORC using the column definitions from a Glue Data Catalog table.
  • B. A separate Glue job rewriting objects afterwards is exactly the extra pipeline the built-in conversion avoids.
  • C. Dynamic partitioning derives S3 prefixes from record content and is independent of format conversion.
  • D. A Lambda transformation runs custom per-record code before conversion; it is not the Parquet conversion mechanism.
Question 5Data Ingestion and Transformation

A 40-shard Amazon Kinesis Data Streams stream currently has six enhanced fan-out consumers. Finance has flagged the enhanced fan-out charges as the largest line on the streaming bill. Reviewing the consumers, the team finds that only two need latency below 200 milliseconds; the other four write to Amazon S3 and Amazon Redshift on schedules measured in minutes, and two of those four simply copy records unchanged into those stores. The team must cut cost without harming the latency-sensitive consumers.

Which two changes should the team make? (Choose TWO.)

Choose 2.

  1. A.

    Reduce the stream from 40 shards to 10 shards, because fewer shards reduce the enhanced fan-out charge for the two remaining dedicated consumers

  2. B.

    Deregister the four latency-tolerant consumers from enhanced fan-out and have them read the stream as standard polling consumers sharing the 2 MB/s per-shard read quota

  3. C.

    Replace the two consumers that write to Amazon S3 and Amazon Redshift with Amazon Data Firehose streams that read the Kinesis data stream as their source and deliver to those destinations, with no consumer application left to operate

  4. D.

    Move the two latency-sensitive consumers off enhanced fan-out as well, and compensate by raising their GetRecords limit parameter to 10,000 records per call

  5. E.

    Keep all six consumers on enhanced fan-out but shorten the stream's retention period to 24 hours, since enhanced fan-out is billed per retained gigabyte

Show answer

Answer: B, C

Enhanced fan-out is billed per consumer-shard-hour, so removing four consumers that do not need it, and replacing two of them with managed Firehose delivery, both cut the bill directly.

  • A. The 40 shards are sized for producer throughput; cutting them to 10 would throttle producers to save the wrong charge.
  • B. Enhanced fan-out is billed per consumer-shard-hour, so removing four consumers that tolerate minutes of latency cuts most of the charge.
  • C. Firehose streams can read a Kinesis data stream and deliver to S3 and to Redshift with no code, removing two consumers from the shared read quota.
  • D. This breaks the stated latency requirement, and the GetRecords limit caps records per call, not throughput or latency.
  • E. Extended retention is billed separately per shard-hour and is not a component of the enhanced fan-out charge.
Question 6Data Ingestion and Transformation

A Step Functions Task state calls a third-party API that returns intermittent HTTP 503 errors, and also occasionally returns a permanent validation error for malformed input. The workflow currently retries every error identically, so malformed input is retried pointlessly and delays the whole execution.

How should the state's error handling be configured?

  1. A.

    Set a single Retry block matching States.ALL with 10 attempts and a high backoff rate, so both error types eventually succeed or exhaust their retries

  2. B.

    Define separate Retry blocks that match the transient error by name with a backoff rate and several attempts, and let the validation error fall through to a Catch block that routes it to a rejects path

  3. C.

    Remove the Retry blocks and wrap the API call in a Lambda function that implements its own retry logic internally

  4. D.

    Set a single Retry block with MaxAttempts of zero so that no error is retried, and rely on a scheduled job to reprocess everything that failed, so that every failure is handled by one recovery path rather than by per-error policies

Show answer

Answer: B

Retriable and non-retriable errors need different handling: retry the transient error by name with backoff, and catch the permanent one to a rejects path.

  • A. This is the current behaviour: the permanent error is retried ten times with backoff before the execution fails.
  • B. Named retries with backoff for the transient error and a Catch for the permanent one handle each failure appropriately.
  • C. Retry logic inside a function is invisible to the state machine and burns function duration while sleeping.
  • D. This forfeits recovery from transient errors and needs an extra scheduled job to do what a retry policy already does.
Question 7Data Ingestion and Transformation

A production Step Functions state machine is updated several times a month. The team wants new definitions to be released to a small share of executions first, and wants a fast way to move all traffic back to the previous definition if errors rise.

Which capability supports this?

  1. A.

    Two separate state machines with different ARNs, and a Lambda function in front that randomly chooses which ARN to start

  2. B.

    AWS CloudFormation change sets, which apply a definition update gradually across executions over a configurable rollout window

  3. C.

    State machine versions together with an alias that routes a configurable percentage of StartExecution calls between two versions

  4. D.

    Step Functions execution history replay, which reruns failed executions against the previous definition automatically

Show answer

Answer: C

Step Functions supports immutable versions and aliases that split StartExecution traffic by weight, which gives both canary release and instant rollback.

  • A. This rebuilds the feature with two ARNs and a routing function to maintain, and no single production identity.
  • B. Change sets preview and apply a stack update; they do not roll a definition out gradually across executions.
  • C. Immutable versions plus a weighted alias give percentage-based canary release and instant rollback to the prior version.
  • D. Redrive restarts failed executions; it does not automatically rerun them against a previous definition.
Question 8Data Ingestion and Transformation

An AWS Glue Spark job reads an Amazon S3 prefix that now contains 12 million small JSON objects spread across a flat key space with no partition structure. Listing the prefix alone takes over 40 minutes before any processing begins, and the job frequently times out. The team can change how the data is organized and how the job reads it.

Which two changes will most reduce the time spent before processing starts? (Choose TWO.)

Choose 2.

  1. A.

    Switch the worker type from G.1X to G.2X, which doubles the memory available to the driver that performs the listing

  2. B.

    Enable input file grouping on the Glue source with groupFiles set to inPartition and a groupSize, so Spark treats many small objects as fewer larger input splits

  3. C.

    Increase the Glue job timeout from 60 minutes to 480 minutes so the listing phase has time to complete before processing begins

  4. D.

    Change the output format of the job from JSON to Apache Parquet, so that the reader spends less time parsing each of the 12 million input objects

  5. E.

    Reorganize the prefix into date-based partitions such as year=/month=/day= and register them in the AWS Glue Data Catalog so the job reads a catalog table with partition predicates instead of listing a flat prefix

Show answer

Answer: B, E

Partitioning plus catalog predicates removes most of the listing, and input file grouping stops Spark from creating one task per tiny object.

  • A. More driver memory may avert an out-of-memory failure but does not reduce the number of list API calls.
  • B. groupFiles and groupSize collapse millions of tiny objects into far fewer splits, removing the task-planning overhead.
  • C. A longer timeout lets the job endure the listing rather than avoiding it, so every run still pays 40 minutes.
  • D. Output format affects downstream readers; the 12 million JSON inputs are still listed and read the same way.
  • E. Partitioned keys plus catalog predicates replace a full prefix enumeration with a metadata lookup and a few prefixes.
Question 9Data Ingestion and Transformation

A four-hour Spark job on Amazon EMR runs on Spot task nodes to save cost. Occasionally a Spot interruption late in the run forces the whole job to restart from the beginning, and the repeated work has started to outweigh the saving.

Which change preserves most of the Spot saving while limiting the cost of an interruption?

  1. A.

    Configure the instance fleet to request a single Spot instance type with the deepest available capacity pool, so that no interruption can occur, and pin the fleet to the single Availability Zone with the most spare capacity

  2. B.

    Increase the Spot maximum price to the On-Demand rate, which guarantees the instances are never reclaimed during the job

  3. C.

    Split the job into shorter stages that write intermediate output to Amazon S3, so an interruption only costs the current stage rather than the whole run

  4. D.

    Move all task nodes back to On-Demand Instances, which removes interruptions entirely and is therefore the lowest total cost option

Show answer

Answer: C

Checkpointing to durable storage bounds the work an interruption destroys, which is what makes long Spot-backed jobs economical.

  • A. A single instance type draws from one capacity pool, which increases rather than eliminates interruption risk.
  • B. Spot instances are reclaimed when EC2 needs the capacity; a higher maximum price does not prevent interruption.
  • C. Durable per-stage output bounds the work an interruption destroys, keeping the Spot discount on most of the compute.
  • D. On-Demand removes interruptions but forfeits the discount entirely; it is not automatically the cheapest outcome.
Question 10Data Ingestion and Transformation

A Python job ingests data from a partner REST API that allows 10 requests per second per client. During peak hours the job receives HTTP 429 responses, and the current code retries immediately in a tight loop, which makes the 429 rate worse. The partner has asked the team to reduce its request rate.

What should the data engineer implement?

  1. A.

    Retry with exponential backoff and jitter, honouring any Retry-After header the API returns, and cap the client's request rate below the published limit

  2. B.

    Retry immediately but from a larger pool of threads, so that a request that fails on one thread is retried on another while the first thread continues

  3. C.

    Request that the partner raise the per-client limit, and keep the existing tight retry loop in place until the new limit takes effect

  4. D.

    Catch the HTTP 429 response, log it, and skip the affected records so the job completes within its scheduled window

Show answer

Answer: A

Exponential backoff with jitter plus client-side rate limiting is the standard way to consume a rate-limited API without amplifying the throttling.

  • A. Backoff spreads retries out, jitter desynchronises concurrent callers, Retry-After respects the server, and a rate cap avoids the limit.
  • B. More threads retrying immediately increases the offered request rate and makes the throttling worse.
  • C. A higher limit with an unchanged tight retry loop simply saturates the new limit as well.
  • D. Skipping throttled records silently loses data, which an ingestion pipeline must never do.

Keep going with 600 more DEA-C01 questions

Free papers every day, in the real exam formats, with progress by exam domain. Unlock every paper and timed mock exam when you are ready.

DEA-C01 sample questions with answers (10 free) · CertifyCloudx