48 lessons
Statistics · Module 1Module material

Describing Data and Reading Statistical Plots interview questions

27 questions interviewers ask on this module’s topics, from first principles to the follow-ups. Each comes with a model answer.

Distributions: histograms, density and the ECDF

  1. What does a histogram's bar height show, and what does changing the bin width change?●
  2. How is a density histogram different from a count histogram?●●
  3. What is a kernel density estimate, and what does its bandwidth control?●●
  4. Why can a histogram not answer "what fraction took 20 minutes or less" when 20 is a bin edge?●
  5. How do you read the 90th percentile from an empirical CDF, and how do you find it by rank?●●
  6. A target says 90% of trips within a whole number of minutes. The 90th percentile is 27.1833. What limit do you report?●●

Center and spread

  1. When would you report the median rather than the mean?●
  2. Removing the largest invoice lowers the mean by £8.41. Does that invoice explain why the mean is above the median?●●
  3. Why do you have to decide the unit of analysis before computing a mean?●
  4. What is the mode, and why is it weak for continuous amounts?●
  5. Why does the sample variance divide by n - 1, and what is its unit?●●
  6. Two groups have almost equal means. The first has the larger standard deviation and the second the larger IQR. How can both be true?●●
  7. One record changes a standard deviation a lot. Should you delete it?●●

Comparing groups

  1. How are box plot whiskers and flagged points defined?●
  2. A point is beyond the whisker. Is it an error? A value is inside the fences. Is it valid?●●
  3. What does a violin plot add to a box plot, and what can it mislead about?●●

Association

  1. What does a Pearson correlation measure, and what are its limits?●
  2. An overall correlation is 0.069. Could the relationship be stronger, or reversed, inside subgroups?●●
  3. What is a confounder, and how would you spot a potential one in exploratory data?●●
  4. Why is an association in observational data not evidence of a causal effect?●
  5. A distance–fare correlation is 0.939. Is fare a function of distance?●●
  6. How do you plot and summarize three million paired points?●●

Categories and who is in the data

  1. University clients have the most defaults. Do they have the highest default rate?●
  2. A dataset has category codes the documentation does not define. What do you do?●●
  3. A survey's raw rows are 39.82% aged 0-19. Can you report that as the national share?●●
  4. Does a random train/test split make a test set representative of the target population?●●
  5. Why compute a correlation heatmap on training rows only, and leave the weight and label out?●●

Directory listings

  • SchoolWhool on Siteefy