ADP sample questions with answers

10 free practice questions for the Associate Data Practitioner exam. Try each one, then open the answer to see why the right option wins and every other option loses.

Question 1Data Preparation and Ingestion

Alderquay Media wants one BigQuery dataset carrying both its Google Ads performance data and the daily CSV extracts its billing system writes to a Cloud Storage bucket under a dated prefix. Both feeds must refresh every morning with no command run by hand, using Google-managed transfers with visible run history rather than the data team's scripts. What should you do?

  1. A.

    Create two BigQuery Data Transfer Service configurations into the dataset: one for the Google Ads source, and one for Cloud Storage using a wildcard URI, each on a daily schedule.

  2. B.

    Create a BigQuery Data Transfer Service configuration for the Google Ads source, and use Database Migration Service to replicate the billing system's CSV extracts into the dataset.

  3. C.

    Deploy a Cloud Run job that calls the Google Ads API and streams rows in with the Storage Write API, and create a Storage Transfer Service job whose sink is the BigQuery dataset.

  4. D.

    Create one Storage Transfer Service job for both feeds, with the BigQuery dataset as its sink, so that the Google Ads export and the daily CSV extracts land together.

Show answer

Answer: A

BigQuery Data Transfer Service has managed connectors for both Google Ads and Cloud Storage, so two scheduled transfer configurations cover both feeds with no code.

  • A. Google Ads and Cloud Storage are both first-party BigQuery Data Transfer Service connectors, and the Cloud Storage one accepts a wildcard URI, so two scheduled configurations cover both feeds with per-run history and no code.
  • B. Database Migration Service migrates relational database instances; it has no notion of loading CSV files from a bucket into BigQuery.
  • C. A custom Cloud Run job reimplements an existing managed connector, and Storage Transfer Service sinks are object stores, so it cannot load rows into a BigQuery table at all.
  • D. Storage Transfer Service moves objects between storage systems and cannot write into BigQuery, and it has no Google Ads source at all.
Question 2Data Preparation and Ingestion

Marrowfield Geophysics holds 480 TB of seismic survey files on an on-premises NAS in a remote field office. The office has one 200 Mbps internet link that survey crews rely on during working hours, and operations will not let it be saturated. The whole archive must reach a Cloud Storage bucket within four weeks. What should you do?

  1. A.

    Install Storage Transfer Service agents on the NAS and run an agent-based transfer job over the existing 200 Mbps link.

  2. B.

    Run gcloud storage rsync with high parallelism from a jump host in the field office to the bucket.

  3. C.

    Order a Transfer Appliance, copy the archive onto it on site, and ship it to Google for upload into the bucket.

  4. D.

    Create a BigQuery Data Transfer Service job that pulls the NAS files into BigQuery and export them to Cloud Storage.

Show answer

Answer: C

480 TB cannot cross a shared 200 Mbps link in four weeks, so the offline Transfer Appliance is the only method that meets the deadline.

  • A. Agent-based Storage Transfer Service is the right online tool for on-premises file systems, but it is still limited by the same 200 Mbps link and misses the deadline by months.
  • B. Parallelism cannot create bandwidth that does not exist, and saturating the link is explicitly forbidden by the operations team.
  • C. Offline shipment bypasses the constrained link entirely and is the supported option for petabyte-scale and multi-hundred-terabyte moves on thin connectivity.
  • D. BigQuery Data Transfer Service loads data into BigQuery from SaaS sources and Cloud Storage; it cannot read an on-premises NAS.
Question 3Data Preparation and Ingestion

Pellamere Clinics runs an 800 GB on-premises MySQL 8 database behind its order entry application. It must move to Cloud SQL for MySQL with minimal downtime: the business accepts a short cutover window, not the hours a full dump and import would take. You want a Google-managed migration that keeps replicating until you promote the destination. What should you do?

  1. A.

    Use Storage Transfer Service to copy the MySQL data directory into Cloud Storage, then attach the copied files to the new Cloud SQL instance as its storage.

  2. B.

    Create a Database Migration Service continuous migration job to Cloud SQL for MySQL, let it replicate changes, and promote the destination at cutover.

  3. C.

    Create a Datastream stream from the MySQL source and write the change events into the Cloud SQL instance.

  4. D.

    Run mysqldump to a file, upload it to Cloud Storage, and import it into the Cloud SQL instance during a maintenance window.

Show answer

Answer: B

Database Migration Service performs a managed initial dump plus continuous change replication into Cloud SQL, so downtime is limited to the promotion step.

  • A. Cloud SQL is a managed service with no mechanism for attaching a raw MySQL data directory copied from on-premises storage.
  • B. DMS continuous migration does the initial load and ongoing replication, leaving only a brief promotion window of downtime.
  • C. Datastream replicates changes mainly to analytical destinations such as BigQuery and Cloud Storage; reaching Cloud SQL needs an extra Dataflow template, and it has no managed promotion step for a migration cutover.
  • D. A one-shot dump and import forces the application offline for the entire load of 800 GB, which the business has rejected.
Question 4Data Preparation and Ingestion

A partner publishes a daily clickstream export of roughly 120 GB into an Amazon S3 bucket. Verroca Athletics needs those objects copied into a Cloud Storage bucket every morning, copying only new or changed objects, with the schedule and run history managed by Google Cloud and no custom code. What should you do?

  1. A.

    Write a Dataflow streaming pipeline that reads the S3 bucket and writes objects to Cloud Storage.

  2. B.

    Create a Storage Transfer Service job with Amazon S3 as the source, the Cloud Storage bucket as the sink, and a daily recurring schedule.

  3. C.

    Order a Transfer Appliance and repeat the shipment cycle each week to pick up the new objects that the partner has written into the S3 bucket.

  4. D.

    Run gcloud storage rsync from a Compute Engine VM started by a cron job every morning.

Show answer

Answer: B

Storage Transfer Service is the managed, scheduled, incremental way to move objects from Amazon S3 into Cloud Storage with no code to maintain.

  • A. Dataflow is a data processing service; using it purely to copy unchanged objects adds code, workers and cost for no transformation benefit.
  • B. A scheduled Storage Transfer Service job copies only new or changed objects from S3 and is fully managed, with built-in run history.
  • C. Transfer Appliance is an offline, one-time bulk method; shipping hardware weekly for a 120 GB daily delta is absurd operationally and far too slow.
  • D. A VM plus cron plus rsync is custom infrastructure the team must patch, monitor and pay for, which the stem rules out.
Question 5Data Preparation and Ingestion

Ostara AI trains models on GPU virtual machines in a single zone, us-central1-a, reading training shards and writing checkpoints continuously. The team needs the lowest possible latency and highest I/O, and accepts that data is redundant only within that zone because the source data is also kept elsewhere. Which Cloud Storage location type fits?

  1. A.

    A regional bucket in us-central1 using the Standard storage class

  2. B.

    A zonal bucket in us-central1-a, created with Rapid Bucket

  3. C.

    A dual-region bucket spanning us-central1 and us-east1

  4. D.

    A multi-region bucket in the US, with turbo replication enabled

Show answer

Answer: B

For AI workloads that need maximum I/O next to compute in one zone, use a zonal bucket through Rapid Bucket.

  • A. A region spreads data across zones, which is more resilient but does not give zonal co-location for maximum I/O.
  • B. Zonal buckets use Rapid storage in the same zone as compute, for the lowest latency and highest I/O.
  • C. Replicating across regions adds cost and distance, which is the opposite of the requirement.
  • D. Turbo replication applies only to dual-region buckets, and a multi-region is not co-located with the GPUs.
Question 6Data Preparation and Ingestion

Draycott Farms loads a partner's daily CSV into an existing BigQuery table with bq load. Recently some rows omit the final optional comments column entirely, and other rows carry an extra trailing column the partner added that you do not need. Both cause load errors. Which two flags let the load succeed without editing the files? (Choose two.)

Choose 2.

  1. A.

    --skipleadingrows=1

  2. B.

    --allowjaggedrows=true

  3. C.

    --replace=true

  4. D.

    --ignoreunknownvalues=true

  5. E.

    --allowquotednewlines=true

Show answer

Answer: B, D

--allowjaggedrows tolerates missing trailing columns, and --ignoreunknownvalues drops extra trailing values in CSV loads.

  • A. This skips a header row at the start of the file, not short or long rows throughout it.
  • B. This accepts rows that are missing trailing optional columns and treats the missing values as NULL.
  • C. This overwrites the table's existing data and schema; it does nothing about malformed rows.
  • D. For CSV this loads rows that carry extra column values and ignores the extra columns.
  • E. This permits newline characters inside quoted fields, which is not the problem described.
Question 7Data Preparation and Ingestion

Kestrel Architecture must copy a 60 TB on-premises NFS share to Cloud Storage and then keep it synchronised weekly. The 2 Gbps office link is shared with video calls, so the transfer must never use more than 500 Mbps. The team wants managed retries and run reports rather than scripts. Which two actions should you take? (Choose two.)

Choose 2.

  1. A.

    Set a bandwidth limit of 500 Mbps on the agent pool that the transfer job uses.

  2. B.

    Schedule gcloud storage rsync from a single server with cron and throttle the network interface.

  3. C.

    Order a Transfer Appliance each week to carry the changed files to Google Cloud.

  4. D.

    Create a BigQuery Data Transfer Service job that reads the NFS share on a weekly schedule.

  5. E.

    Install Storage Transfer Service agents on machines that can read the NFS share and group them in an agent pool.

Show answer

Answer: A, E

Agent-based Storage Transfer Service with a bandwidth-limited agent pool gives a managed, recurring, capped transfer from NFS.

  • A. Agent pools provide control over transfer bandwidth limits, which enforces the 500 Mbps cap.
  • B. This is the script-based approach the team wants to avoid, with no managed retries or run reports.
  • C. The link can move 60 TB in weeks and the weekly change is far smaller, so shipping hardware each week is needlessly slow and laborious.
  • D. BigQuery Data Transfer Service loads into BigQuery and cannot read an on-premises file share.
  • E. File system sources need Storage Transfer Service agents in an agent pool; the service then manages scheduling, retries and reporting.
Question 8Data Preparation and Ingestion

Selwyn Logistics' shipment table in BigQuery records destination states inconsistently, for example 'CA', 'Calif.', 'california' and 'Cal'. There are dozens of variants across all states, and operations staff add new ones as they discover them. You need a standardised state code that non-engineers can maintain. What should you do?

  1. A.

    Apply UPPER(state) to the column, which converts all of the variants to the same state code.

  2. B.

    Keep a mapping table of variant-to-code pairs that operations maintains, and join it on LOWER(TRIM(state)) during the ELT step.

  3. C.

    Write a CASE expression listing every variant in the transformation SQL and edit it when new ones appear.

  4. D.

    Delete the rows whose state value is not already a valid two-letter code.

Show answer

Answer: B

Standardise many free-text variants with a maintained mapping table joined on a normalised key.

  • A. Upper-casing turns 'Calif.' into 'CALIF.', not 'CA', so most variants remain distinct.
  • B. A maintained reference table handles any number of variants, and staff can add rows without changing SQL.
  • C. It works initially, but every new variant needs an engineer to change and redeploy the SQL.
  • D. Deleting real shipments loses data that can be recovered by mapping, which is the purpose of cleaning.
Question 9Data Preparation and Ingestion

Linnet Publishing serves e-book cover images and sample chapters to readers across North America from Cloud Storage. It wants the data to survive a regional outage with no change of storage paths, does not need to choose specific regions, and prefers lower storage cost than a dual-region. Which location type should you choose?

  1. A.

    A multi-region bucket, such as the US multi-region

  2. B.

    A dual-region bucket, with specific regions chosen

  3. C.

    A regional bucket, such as one in us-east1

  4. D.

    A zonal bucket in one us-east1 zone, using Rapid Bucket

Show answer

Answer: A

Multi-region buckets give cross-region redundancy without choosing regions, at lower storage cost than dual-regions, and suit content serving.

  • A. Multi-regions provide cross-region redundancy at a lower storage price than dual-regions and suit content serving.
  • B. Dual-regions give precise control of regions at a higher storage price, which the stem says is not needed.
  • C. A single region does not survive the loss of that region.
  • D. A zonal bucket keeps data in one zone for maximum performance and has no cross-region redundancy.
Question 10Data Preparation and Ingestion

Stellan Airways uses Dataform to build a curated bookings table with a uniqueKey assertion on booking_id, followed by a revenue summary table that references it. Last night the assertion failed, yet the revenue summary still rebuilt from the duplicated rows and fed finance reports. How should you prevent this from recurring?

  1. A.

    Replace the assertion with a nonNull assertion, because Dataform halts downstream actions only when a nonNull assertion fails.

  2. B.

    Increase the BigQuery reservation so that the assertion finishes before the revenue summary starts running.

  3. C.

    Move the uniqueKey assertion into the revenue summary's config block so it runs before that table is built.

  4. D.

    Make the bookings table's assertions dependencies of the revenue summary, for example with includeDependentAssertions.

Show answer

Answer: D

In Dataform, failing assertions only block downstream actions when the assertions are declared as dependencies of those actions.

  • A. No assertion type halts dependents automatically; the fix is the dependency setting, not the assertion type.
  • B. The problem is dependency configuration, not timing; a faster assertion would still not stop the dependent action.
  • C. Assertions in a table's config block run after that table is created, so the summary would still build first.
  • D. Dataform does not block a dependent action when a dependency's assertions fail unless those assertions are set as dependencies.

Keep going with 490 more ADP questions

Free papers every day, in the real exam formats, with progress by exam domain. Unlock every paper and timed mock exam when you are ready.

ADP sample questions with answers (10 free) · CertifyCloudx