Professional Data Engineer Exam - Question 235

Question

You want to schedule a number of sequential load and transformation jobs. Data files will be added to a Cloud Storage bucket by an upstream process. There is no fixed schedule for when the new data arrives. Next, a Dataproc job is triggered to perform some transformations and write the data to BigQuery. You then need to run additional transformation jobs in BigQuery. The transformation jobs are different for every table. These jobs might take hours to complete. You need to determine the most efficient and maintainable workflow to process hundreds of tables and provide the freshest data to your end users. What should you do?

Examice · Accepted Answer

To schedule and manage the sequential load and transformation jobs in an efficient and maintainable way, creating a single Apache Airflow DAG that handles all tables within the pipeline is ideal. Given that new data files can arrive at any time, using Cloud Storage object triggers to launch a Cloud Function, which then triggers the DAG, ensures the workflow begins processing immediately when new data is available. By using Dataproc and BigQuery operators, you can handle the necessary transformations efficiently. This approach minimizes complexity and maintenance, as it avoids the need to manage separate DAGs for each table, making it more scalable and easier to maintain.

cuadradobertolinisebastiancami · Answer

D

* Transformations are in Dataproc and BigQuery. So you don't need operators for GCS (A and B can be discard)
* "There is no fixed schedule for when the new data arrives." so you trigger the DAG when a file arrives
* "The transformation jobs are different for every table. " so you need a DAG for each table.

Then, D is the most suitable answer

Jordan18 · Answer

why not C?

Matt_108 · Answer

Option D, which gets triggered when the data comes in and accounts for the fact that each table has its own set of transformations

raaad · Answer

- Option D: Tailored handling and scheduling for each table; triggered by data arrival for more timely and efficient processing.

scaenruy · Answer

D. 
1. Create an Apache Airflow directed acyclic graph (DAG) in Cloud Composer with sequential tasks by using the Dataproc and BigQuery operators.
2. Create a separate DAG for each table that needs to go through the pipeline.
3. Use a Cloud Storage object trigger to launch a Cloud Function that triggers the DAG.

JyoGCP · Answer

Option D

8ad5266 · Answer

This explains why it's not D:
maintainable workflow to process hundreds of tables and provide the freshest data to your end users

How is creating a DAG for each of the hundreds of tables maintainable?

Professional Data Engineer Exam - Question 235

Discussion