An ETL pipeline extracts from an operational database, applies three successive transformation stages, and loads the result into Amazon Redshift. Each stage is a separate AWS Glue job, and a failed stage must be re-runnable without re-extracting from the operational database.
Where should the output of each stage be written?
- A.
To a stage-specific Amazon S3 prefix, so each job reads the previous stage's output and can be re-run independently
- B.
To the local disk of the Glue workers, so intermediate results avoid Amazon S3 request charges and stay close to the compute
- C.
To a Spark in-memory DataFrame passed directly between the three jobs, avoiding the cost of writing intermediate results to storage
- D.
Back to the operational database in temporary tables, so all intermediate state is held in one transactional system
Show answer
Answer: A
Durable, addressable staging in Amazon S3 is what makes each stage independently re-runnable without touching the source.
- A. Durable, addressable per-stage output survives the job, supports inspection, and makes each stage independently re-runnable.
- B. Glue worker local disk is deallocated when the run ends, so the next stage finds nothing and a failed stage cannot resume.
- C. Separate Glue jobs are separate Spark applications; in-memory DataFrames cannot cross between them and do not survive failure.
- D. This puts ETL load and storage on the operational system and couples pipeline reliability to a transactional database.
