Summary of Understanding Statistical Fallacies and Misrepresentation

Understanding Statistical Fallacies and Misrepresentation

Introduction

Statistics are powerful tools for understanding the world, but they can be used (intentionally or not) to mislead. This guide breaks down common techniques that distort data so you can spot them, ask the right questions, and make better judgments.

Definition: A misleading statistic is a numeric claim or visual that gives an incorrect impression about reality because of biased samples, inappropriate summaries, missing context, or deceptive presentation.

1. The Sample with the Built-in Bias

What it means

A sample is biased when it does not fairly represent the population you want to describe. Bias can come from how people are selected or from who chooses to respond.

Definition: A biased sample is one in which some members of the intended population are systematically more likely to be included than others.

How bias arises

  • Nonrandom selection of participants
  • Self-selection: people who respond differ from those who do not
  • Excluding groups unintentionally

Practical example

A university surveys graduates about salary. If high earners are more likely to reply, the reported average salary will be higher than the true average.

Questions to ask

  • Who was surveyed?
  • How were they chosen?
  • How many responded?
  • Is any group missing or underrepresented?
💡 Did you know?Fun fact: People with stronger opinions or experiences are more likely to respond to surveys, which is why polls often show polarized views.

2. The Well-Chosen Average

The three averages

  • Mean: the sum of values divided by the number of values, e.g., $\text{mean} = \frac{\sum x_i}{n}$.
  • Median: the middle value when data are sorted; robust to extreme values.
  • Mode: the most common value in a dataset.

Definition: The mean measures central tendency by arithmetic average, the median is the 50% cutoff, and the mode is the most frequent observation.

Why the choice matters

Different averages can give very different impressions when data are skewed or contain outliers.

Example

In a group where nine people earn $20{,}000$ and one person earns $1{,}000{,}000$, the mean is much higher than what most people earn; the median better reflects the typical income.

Quick checklist

  • Ask which average is used.
  • Check for outliers that affect the mean.
  • Prefer median for skewed distributions.
💡 Did you know?Did you know that the median income often gives a better sense of a typical household than the mean when income distribution is skewed?

3. The Little Figures That Are Not There

Missing context and small samples

Small samples or omitted details can make results look impressive even when they are unreliable.

Definition: Sample size is the number of observations; small samples have high variability and give weaker evidence.

Why small samples mislead

  • They have large random variation
  • Results may not replicate with more data
  • Percentages from small counts appear dramatic (e.g., "50% improvement" from 1 to 2 cases)

Example

A medicine tested on 3 people that helps 2 looks like a 67% success rate, but the sample is too small to trust that percentage.

What to check

  • How many subjects were involved?
  • Are confidence intervals or error estimates provided?
  • Were any important details omitted (methods, exclusions)?
💡 Did you know?Fun fact: Random chance can produce surprising-looking results in small samples; gamblers’ runs often look meaningful though they are just luck.

4. Much Ado about Practically Nothing

Statistical vs. practical significance

A result can be statistically detectable yet too small to matter in practice.

Definition: Statistical significance indicates an effect likely not due to random chance; practical significance assesses whether the effect size matters in the real world.

Margin of error and random variation

  • All estimates have uncertainty (margin of error).
  • Small differences inside the margin of error may be meaningless.

Example

Product A preferred by $51%$ and product B b

Sign up for the full summary
FlashcardsKnowledge testSummaryPodcastMindmap
Start for free

Already have an account? Sign in

Statistical Misleading Techniques

Klíčové pojmy: Ask who was surveyed and how participants were selected, Check sample size and response rate before trusting results, Distinguish mean, median, and mode; choose median for skewed data, Watch for outliers that inflate the mean, Be skeptical of percentages from very small samples, Ask for margin of error or confidence intervals to assess uncertainty, Evaluate whether differences are practically meaningful, not just statistically significant, Inspect graphs for axis truncation, scale manipulation, and missing labels, Demand missing details: methods, exclusions, and how data were collected, Prefer replication or larger studies when initial samples are small, Question self-selected surveys and polls for response bias, Compare numerical values behind visuals rather than relying on graphics

## Introduction Statistics are powerful tools for understanding the world, but they can be used (intentionally or not) to mislead. This guide breaks down common techniques that distort data so you can spot them, ask the right questions, and make better judgments. > Definition: A misleading statistic is a numeric claim or visual that gives an incorrect impression about reality because of biased samples, inappropriate summaries, missing context, or deceptive presentation. ## 1. The Sample with the Built-in Bias ### What it means A sample is biased when it does not fairly represent the population you want to describe. Bias can come from how people are selected or from who chooses to respond. > Definition: A biased sample is one in which some members of the intended population are systematically more likely to be included than others. ### How bias arises - Nonrandom selection of participants - Self-selection: people who respond differ from those who do not - Excluding groups unintentionally ### Practical example A university surveys graduates about salary. If high earners are more likely to reply, the reported average salary will be higher than the true average. ### Questions to ask - Who was surveyed? - How were they chosen? - How many responded? - Is any group missing or underrepresented? Fun fact: People with stronger opinions or experiences are more likely to respond to surveys, which is why polls often show polarized views. ## 2. The Well-Chosen Average ### The three averages - **Mean**: the sum of values divided by the number of values, e.g., $\text{mean} = \frac{\sum x_i}{n}$. - **Median**: the middle value when data are sorted; robust to extreme values. - **Mode**: the most common value in a dataset. > Definition: The mean measures central tendency by arithmetic average, the median is the 50% cutoff, and the mode is the most frequent observation. ### Why the choice matters Different averages can give very different impressions when data are skewed or contain outliers. ### Example In a group where nine people earn $20{,}000$ and one person earns $1{,}000{,}000$, the mean is much higher than what most people earn; the median better reflects the typical income. ### Quick checklist - Ask which average is used. - Check for outliers that affect the mean. - Prefer median for skewed distributions. Did you know that the median income often gives a better sense of a typical household than the mean when income distribution is skewed? ## 3. The Little Figures That Are Not There ### Missing context and small samples Small samples or omitted details can make results look impressive even when they are unreliable. > Definition: Sample size is the number of observations; small samples have high variability and give weaker evidence. ### Why small samples mislead - They have large random variation - Results may not replicate with more data - Percentages from small counts appear dramatic (e.g., "50% improvement" from 1 to 2 cases) ### Example A medicine tested on 3 people that helps 2 looks like a 67% success rate, but the sample is too small to trust that percentage. ### What to check - How many subjects were involved? - Are confidence intervals or error estimates provided? - Were any important details omitted (methods, exclusions)? Fun fact: Random chance can produce surprising-looking results in small samples; gamblers’ runs often look meaningful though they are just luck. ## 4. Much Ado about Practically Nothing ### Statistical vs. practical significance A result can be statistically detectable yet too small to matter in practice. > Definition: Statistical significance indicates an effect likely not due to random chance; practical significance assesses whether the effect size matters in the real world. ### Margin of error and random variation - All estimates have uncertainty (margin of error). - Small differences inside the margin of error may be meaningless. ### Example Product A preferred by $51\%$ and product B b