Skip to main content
๐Ÿ“˜ Complete PDF

PySpark Certification Practice Test 07

Find The Problem In The Code

Question 1 of 20
๐Ÿ“‹ Multiple Choice
Question 121
A developer writes: ```python df = spark.read.parquet("/sales") result = df.collect() print(len(result)) ``` The dataset contains 2 billion records. What is the biggest problem?
A. Missing Cache
โœ“
B. Driver Out Of Memory Risk
C. Missing Repartition
D. Missing Broadcast
Answer
โœ… B. Driver Out Of Memory Risk
Explanation
### Explanation collect() moves all data to the Driver.
Answered: 0 / 20
Question Navigator (Test 07)0% Answered