Login to manage your account

Please enter a valid email address.
Forgot Password?
Please enter a valid password.
OR

Don't have an account yet? Sign up

What is data skew in a distributed job?

  1. A Uneven distribution of data across partitions, so some tasks take far longer than others
  2. B Corrupted records
  3. C Incorrect schema inference
  4. D Duplicate partitions
Answer

Uneven distribution of data across partitions, so some tasks take far longer than others

A job runs only as fast as its slowest task, so a single hot key can dominate total runtime regardless of cluster size.

All Big data MCQs

Login to manage your account

Please enter a valid email address.
Forgot Password?
Please enter a valid password.
OR

Don't have an account yet? Sign up as