Statistics · Module 1Module material
Describing Data and Reading Statistical Plots interview questions
27 questions interviewers ask on this module’s topics, from first principles to the follow-ups. Each comes with a model answer.
Distributions: histograms, density and the ECDF
- What does a histogram's bar height show, and what does changing the bin width change?●
- How is a density histogram different from a count histogram?●●
- What is a kernel density estimate, and what does its bandwidth control?●●
- Why can a histogram not answer "what fraction took 20 minutes or less" when 20 is a bin edge?●
- How do you read the 90th percentile from an empirical CDF, and how do you find it by rank?●●
- A target says 90% of trips within a whole number of minutes. The 90th percentile is
27.1833. What limit do you report?●●
Center and spread
- When would you report the median rather than the mean?●
- Removing the largest invoice lowers the mean by
£8.41. Does that invoice explain why the mean is above the median?●● - Why do you have to decide the unit of analysis before computing a mean?●
- What is the mode, and why is it weak for continuous amounts?●
- Why does the sample variance divide by
n - 1, and what is its unit?●● - Two groups have almost equal means. The first has the larger standard deviation and the second the larger IQR. How can both be true?●●
- One record changes a standard deviation a lot. Should you delete it?●●
Comparing groups
- How are box plot whiskers and flagged points defined?●
- A point is beyond the whisker. Is it an error? A value is inside the fences. Is it valid?●●
- What does a violin plot add to a box plot, and what can it mislead about?●●
Association
- What does a Pearson correlation measure, and what are its limits?●
- An overall correlation is
0.069. Could the relationship be stronger, or reversed, inside subgroups?●● - What is a confounder, and how would you spot a potential one in exploratory data?●●
- Why is an association in observational data not evidence of a causal effect?●
- A distance–fare correlation is
0.939. Is fare a function of distance?●● - How do you plot and summarize three million paired points?●●
Categories and who is in the data
- University clients have the most defaults. Do they have the highest default rate?●
- A dataset has category codes the documentation does not define. What do you do?●●
- A survey's raw rows are
39.82%aged0-19. Can you report that as the national share?●● - Does a random train/test split make a test set representative of the target population?●●
- Why compute a correlation heatmap on training rows only, and leave the weight and label out?●●