Summary of Understanding Statistical Fallacies and Misrepresentation
Understanding Statistical Fallacies and Misrepresentation
Introduction
Statistics are powerful tools for understanding the world, but they can be used (intentionally or not) to mislead. This guide breaks down common techniques that distort data so you can spot them, ask the right questions, and make better judgments.
Definition: A misleading statistic is a numeric claim or visual that gives an incorrect impression about reality because of biased samples, inappropriate summaries, missing context, or deceptive presentation.
1. The Sample with the Built-in Bias
What it means
A sample is biased when it does not fairly represent the population you want to describe. Bias can come from how people are selected or from who chooses to respond.
Definition: A biased sample is one in which some members of the intended population are systematically more likely to be included than others.
How bias arises
- Nonrandom selection of participants
- Self-selection: people who respond differ from those who do not
- Excluding groups unintentionally
Practical example
A university surveys graduates about salary. If high earners are more likely to reply, the reported average salary will be higher than the true average.
Questions to ask
- Who was surveyed?
- How were they chosen?
- How many responded?
- Is any group missing or underrepresented?
2. The Well-Chosen Average
The three averages
- Mean: the sum of values divided by the number of values, e.g., $\text{mean} = \frac{\sum x_i}{n}$.
- Median: the middle value when data are sorted; robust to extreme values.
- Mode: the most common value in a dataset.
Definition: The mean measures central tendency by arithmetic average, the median is the 50% cutoff, and the mode is the most frequent observation.
Why the choice matters
Different averages can give very different impressions when data are skewed or contain outliers.
Example
In a group where nine people earn $20{,}000$ and one person earns $1{,}000{,}000$, the mean is much higher than what most people earn; the median better reflects the typical income.
Quick checklist
- Ask which average is used.
- Check for outliers that affect the mean.
- Prefer median for skewed distributions.
3. The Little Figures That Are Not There
Missing context and small samples
Small samples or omitted details can make results look impressive even when they are unreliable.
Definition: Sample size is the number of observations; small samples have high variability and give weaker evidence.
Why small samples mislead
- They have large random variation
- Results may not replicate with more data
- Percentages from small counts appear dramatic (e.g., "50% improvement" from 1 to 2 cases)
Example
A medicine tested on 3 people that helps 2 looks like a 67% success rate, but the sample is too small to trust that percentage.
What to check
- How many subjects were involved?
- Are confidence intervals or error estimates provided?
- Were any important details omitted (methods, exclusions)?
4. Much Ado about Practically Nothing
Statistical vs. practical significance
A result can be statistically detectable yet too small to matter in practice.
Definition: Statistical significance indicates an effect likely not due to random chance; practical significance assesses whether the effect size matters in the real world.
Margin of error and random variation
- All estimates have uncertainty (margin of error).
- Small differences inside the margin of error may be meaningless.
Example
Product A preferred by $51%$ and product B b
Already have an account? Sign in
Statistical Misleading Techniques
Klíčové pojmy: Ask who was surveyed and how participants were selected, Check sample size and response rate before trusting results, Distinguish mean, median, and mode; choose median for skewed data, Watch for outliers that inflate the mean, Be skeptical of percentages from very small samples, Ask for margin of error or confidence intervals to assess uncertainty, Evaluate whether differences are practically meaningful, not just statistically significant, Inspect graphs for axis truncation, scale manipulation, and missing labels, Demand missing details: methods, exclusions, and how data were collected, Prefer replication or larger studies when initial samples are small, Question self-selected surveys and polls for response bias, Compare numerical values behind visuals rather than relying on graphics