AWSAssociate

AWS Certified Data Engineer – Associate

DEA-C01

Implement data pipelines, manage data stores, and ensure data quality, security and governance on AWS.

Duration
130 min
Exam questions
65
Passing score
720 / 1000
Exam fee
$150
Question formats:Multiple choiceMultiple response
Free plan
3 free papers
Free account
Mocks locked
Pro only
Upgrade to Pro
Every paper and mock exam.
See Pro

Start with free papers Free

You get 3 free practice papers with your plan.

Free
Mixed paper 1
Domain 1 · 25 questions
Free
Mixed paper 2
Domain 1 · 25 questions
Free
Mixed paper 3
Domain 1 · 25 questions

Domain papers 610 questions

Free
Mixed paper 1
25 questions · 60 min
Free
Mixed paper 2
25 questions · 60 min
Free
Mixed paper 3
25 questions · 60 min
Pro
Mixed paper 4
25 questions · 60 min
Pro
Mixed paper 5
25 questions · 60 min
Pro
Mixed paper 6
25 questions · 60 min
Pro
Mixed paper 7
21 questions · 51 min

Mock exams Pro

Full-length, exam-like practice tests. Available with Pro.

Mock exam 1
65 questions · 130 min
Mock exam 2
64 questions · 130 min
Mock exam 3
65 questions · 130 min

Try a sample question

All 10 sample questions →
Question 1Data Ingestion and Transformation

An ETL pipeline extracts from an operational database, applies three successive transformation stages, and loads the result into Amazon Redshift. Each stage is a separate AWS Glue job, and a failed stage must be re-runnable without re-extracting from the operational database.

Where should the output of each stage be written?

  1. A.

    To a stage-specific Amazon S3 prefix, so each job reads the previous stage's output and can be re-run independently

  2. B.

    To the local disk of the Glue workers, so intermediate results avoid Amazon S3 request charges and stay close to the compute

  3. C.

    To a Spark in-memory DataFrame passed directly between the three jobs, avoiding the cost of writing intermediate results to storage

  4. D.

    Back to the operational database in temporary tables, so all intermediate state is held in one transactional system

Show answer

Answer: A

Durable, addressable staging in Amazon S3 is what makes each stage independently re-runnable without touching the source.

  • A. Durable, addressable per-stage output survives the job, supports inspection, and makes each stage independently re-runnable.
  • B. Glue worker local disk is deallocated when the run ends, so the next stage finds nothing and a failed stage cannot resume.
  • C. Separate Glue jobs are separate Spark applications; in-memory DataFrames cannot cross between them and do not survive failure.
  • D. This puts ETL load and storage on the operational system and couples pipeline reliability to a transactional database.

What's on the exam

4 domains · 17 task statements, straight from the official exam guide (as of 2026-09-29).

  1. 1.1Perform data ingestion
    • Skill 1.1.1: Read data from streaming sources (for example, Amazon Kinesis, Amazon Managed Streaming for Apache Kafka [Amazon MSK], Amazon DynamoDB Streams, AWS DMS, AWS Glue, Amazon Redshift).
    • Skill 1.1.2: Read data from batch sources (for example, Amazon S3, AWS Glue, Amazon EMR, AWS DMS, Amazon Redshift, AWS Lambda, Amazon AppFlow).
    • Skill 1.1.3: Implement appropriate configuration options for batch ingestion.
    • Skill 1.1.4: Consume data APIs.
    • Skill 1.1.5: Set up schedulers by using Amazon EventBridge, Apache Airflow, or time-based schedules for jobs and crawlers.
    • Skill 1.1.6: Set up event triggers (for example, Amazon S3 Event Notifications, EventBridge).
    • Skill 1.1.7: Call a Lambda function from Kinesis.
    • Skill 1.1.8: Create allowlists for IP addresses to allow connections to data sources.
    • Skill 1.1.9: Implement throttling and overcoming rate limits (for example, DynamoDB, Amazon RDS, Kinesis).
    • Skill 1.1.10: Manage fan-in and fan-out for streaming data distribution.
    • Skill 1.1.11: Describe replayability of data ingestion pipelines.
    • Skill 1.1.12: Define stateful and stateless data transactions.
  2. 1.2Transform and process data
    • Skill 1.2.1: Optimize container usage for performance needs (for example, Amazon EKS, Amazon ECS).
    • Skill 1.2.2: Connect to different data sources (for example, Java Database Connectivity [JDBC], Open Database Connectivity [ODBC]).
    • Skill 1.2.3: Integrate data from multiple sources.
    • Skill 1.2.4: Optimize costs while processing data.
    • Skill 1.2.5: Implement data transformation services based on requirements (for example, Amazon EMR, AWS Glue, Lambda, Amazon Redshift).
    • Skill 1.2.6: Transform data between formats (for example, from .csv to Apache Parquet).
    • Skill 1.2.7: Troubleshoot and debug common transformation failures and performance issues.
    • Skill 1.2.8: Create data APIs to make data available to other systems by using AWS services.
    • Skill 1.2.9: Define volume, velocity, and variety of data (for example, structured data, unstructured data).
    • Skill 1.2.10: Integrate large language models (LLMs) for data processing.
  3. 1.3Orchestrate data pipelines
    • Skill 1.3.1: Use orchestration services to build workflows for data ETL pipelines (for example, Lambda, EventBridge, Amazon Managed Workflows for Apache Airflow [Amazon MWAA], AWS Step Functions, AWS Glue workflows).
    • Skill 1.3.2: Build data pipelines for performance, availability, scalability, resiliency, and fault tolerance.
    • Skill 1.3.3: Implement and maintain serverless workflows.
    • Skill 1.3.4: Use notification services to send alerts (for example, Amazon SNS, Amazon SQS).
  4. 1.4Apply programming concepts
    • Skill 1.4.1: Optimize code to reduce runtime for data ingestion and transformation.
    • Skill 1.4.2: Configure Lambda functions to meet concurrency and performance needs.
    • Skill 1.4.3: Use programming languages and frameworks for data engineering (for example, Python, SQL, Scala, R, Java, Bash, PowerShell).
    • Skill 1.4.4: Use software engineering best practices for data engineering (for example, version control, testing, logging, monitoring).
    • Skill 1.4.5: Use Infrastructure as Code (IaC) to deploy data engineering solutions.
    • Skill 1.4.6: Use AWS SAM to package and deploy serverless data pipelines (for example, Lambda functions, Step Functions, DynamoDB tables).
    • Skill 1.4.7: Use and mount storage volumes from within Lambda functions.
    • Skill 1.4.8: Use infrastructure as code (IaC) for repeatable resource deployment (for example, AWS CloudFormation and AWS CDK).
    • Skill 1.4.9: Describe continuous integration and continuous delivery (CI/CD) (implementation, testing, and deployment of data pipelines).
    • Skill 1.4.10: Define distributed computing.
    • Skill 1.4.11: Describe data structures and algorithms (for example, graph data structures and tree data structures).

Outline reproduced from the vendor's public exam guide for study reference.Official guide

DEA-C01 practice — frequently asked questions

Are these real DEA-C01 exam questions?

No. Every question on CertifyCloudx is original, written by us against Amazon Web Services's publicly available DEA-C01 exam guide to rehearse the skills it lists. None are actual exam questions, and CertifyCloudx is not affiliated with or endorsed by Amazon Web Services.

How many DEA-C01 practice questions are there?

610 practice questions, including 3 full-length timed mock exams and 59 domain papers of up to 25 questions (mixed and by topic). Every question has a detailed explanation of why the right answer wins and why each distractor loses.

Is the content up to date with the current DEA-C01 exam guide?

The questions are written against the DEA-C01 exam guide dated 2026-09-29, and we revise them when Amazon Web Services updates the guide.

What question formats are covered?

The same formats the real DEA-C01 uses: Multiple choice, Multiple response. Each is rendered and graded the way the exam does it.

How long is the DEA-C01 exam and how many questions does it have?

According to Amazon Web Services's published exam details: 65 questions, 130 minutes, passing score 720 / 1000. Our mock exams use the same time limit and question count. Always confirm current details with Amazon Web Services before booking.

Can I practise DEA-C01 for free?

Yes. 3 papers are free, with up to 10 questions a day on the free plan and no card needed. Pro unlocks every paper and mock exam with no daily limit.

Does CertifyCloudx guarantee that I will pass?

No practice material can guarantee a result. CertifyCloudx helps you find and close your weak areas — accuracy by exam-guide domain and topic shows what to study next.