DATABRICKS-CERTIFIED-ASSOCIATE-DEVELOPER-FOR-APACHE-SPARK Web TestEngine demo

Exit VCEDump DATABRICKS-CERTIFIED-ASSOCIATE-DEVELOPER-FOR-APACHE-SPARK Databricks Certified Associate Developer for Apache Spark
Question 51 of 82
0% complete
Q51 Single choice

A data engineer is working with a dataset containing 2 billion rows distributed across 10 Spark partitions. The engineer needs to compute the approximate count of distinct users by their 'user_id' field and also calculate the average value of 'transaction_amount'. Both calculations must be done in a single transformation step to minimize shuffling.

Which set of Spark code will achieve this goal while ensuring performance?

Sign in to mark questions

Sign in to save marked questions and return to this demo.

Sign in